Benchmarks measure what AI can do.
We measure what it actually does.

Error rates, costs, and failure modes from AI systems running in production — ours. Every figure here comes from a run we can point to.

Latest report — August 5, 2026

1 in 3

factual claims had no support in the sources the system was given. Not wild hallucination — plausible context filled in from model memory.

Verified 65.8%Flagged 29.7%Too recent 2.7%Corrected 1.8%

n = 1,140 articles / 5,500 individual claims · window 2026-05-03 → 2026-08-04