Benchmarks measure what AI can do.
We measure what it actually does.
Error rates, costs, and failure modes from AI systems running in production — ours. Every figure here comes from a run we can point to.
1 in 3
factual claims had no support in the sources the system was given. Not wild hallucination — plausible context filled in from model memory.
Verified 65.8%Flagged 29.7%Too recent 2.7%Corrected 1.8%