Independent eval
Independently verified: zero errors, 8x cheaper, on real enterprise data.
Not our benchmark. BoonAI ran OpenSymbolic against their own production system, on their own document corpus, scored by an independent judge.
01
The setup
BoonAI evaluated OpenSymbolic against their production multi-turn RAG system on 104 real enterprise queries over their own construction-document corpus: a mix of factual, multi-hop, document-comparison, and vision questions. Answers were scored by an LLM judge; failures scored 0.1 and were never excluded. Same task, same data, head to head. Recurring, high-volume, document-heavy work.
02
Headline results
03
Per-category breakdown
*The vision dip was an S3 permissions issue during the eval, not a framework limitation.
Biggest gains on multi-hop and doc comparison, the hard, high-value queries where the baseline's context overflows caused silent failures.
Head-to-head: 27 OpenSymbolic wins · 29 baseline · 48 ties
04
Why it's cheaper
The planner sees summaries and highlights, never raw page content. Context grows roughly 0.5 to 2K tokens per step instead of the 5 to 30K a raw-data loop re-sends every iteration. 64% of queries resolve in a single iteration, with zero errors.
05
Public benchmarks
+22.9 pts over the next-best approach across 2,556 multi-document queries. 99.6% goal completion.
Pass rate at 4 to 13x lower cost than a standard tool-calling loop. 7 of 11 models hit 100%.
Near the 96% human bar, backed by a theorem prover rather than a guess.
Want these numbers on your data?
We'll prove it on one of your workflows in two weeks, measured head-to-head against what you run today.
Book a pilot