Independent eval

Independently verified: zero errors, 8x cheaper, on real enterprise data.

Not our benchmark. BoonAI ran OpenSymbolic against their own production system, on their own document corpus, scored by an independent judge.

01

The setup

BoonAI evaluated OpenSymbolic against their production multi-turn RAG system on 104 real enterprise queries over their own construction-document corpus: a mix of factual, multi-hop, document-comparison, and vision questions. Answers were scored by an LLM judge; failures scored 0.1 and were never excluded. Same task, same data, head to head. Recurring, high-volume, document-heavy work.

02

Headline results

MetricBaselineOpenSymbolicChange
Errors13.5%0%zero failed queries
Cost / query$0.61$0.088x cheaper
Tokens (median)67,2629,45086% fewer
Latency (median)46.5s29.7s36% faster
Accuracy (mean)0.7770.813+4.6%

03

Per-category breakdown

CategorynBaselineOpenSymbolicTokens savedLatency
Simple factual550.9240.898 89% fewer9% faster
Multi-hop250.5840.744 (+0.160)80% fewer45% faster
Doc comparison150.5600.655 (+0.095)92% fewer47% faster
Vision selection90.7830.753* 93% fewer38% faster

*The vision dip was an S3 permissions issue during the eval, not a framework limitation.

Biggest gains on multi-hop and doc comparison, the hard, high-value queries where the baseline's context overflows caused silent failures.

Head-to-head: 27 OpenSymbolic wins · 29 baseline · 48 ties

04

Why it's cheaper

The planner sees summaries and highlights, never raw page content. Context grows roughly 0.5 to 2K tokens per step instead of the 5 to 30K a raw-data loop re-sends every iteration. 64% of queries resolve in a single iteration, with zero errors.

0.5 to 2Ktokens per step, OpenSymbolic
5 to 30Ktokens per step, raw-data loop
How it works →

05

Public benchmarks

MultiHop-RAG82.9%

+22.9 pts over the next-best approach across 2,556 multi-document queries. 99.6% goal completion.

TravelPlanner (ICML 2024)97.9%

Pass rate at 4 to 13x lower cost than a standard tool-calling loop. 7 of 11 models hit 100%.

FOLIO (first-order logic)89.2%

Near the 96% human bar, backed by a theorem prover rather than a guess.

Want these numbers on your data?

We'll prove it on one of your workflows in two weeks, measured head-to-head against what you run today.

Book a pilot