Reducing hallucinations in document extraction
How citations and confidence keep extracted values grounded in the source.
How citations and confidence keep extracted values grounded in the source.
The problem
Every team that works with documents runs into grounded extraction eventually. The naive approach gets you 80% of the way and then stalls.
What we did
We treated it as a measurement problem first. Before changing the model, we built a labeled set of the hardest examples we could find, so we could tell whether any change actually helped.
From there, the work was iterative: adjust the model, score it on the hard set, keep what moved the number, and throw out what didn't. No change ships without evidence.
- A benchmark built from real, difficult documents.
- Targeted training on the failure cases.
- Confidence signals so downstream systems know when to check.
Results
Want to see it on your own documents? Open the playground or reach out — we're happy to run a sample with you.