Docen
← BLOG·ENGINEERING

High-fidelity OCR drives accurate structured extraction

Extraction is only as good as the text underneath it. Recognition quality compounds.

THE DOCEN TEAM·Jun 3, 2025·6 MIN READ

Extraction is only as good as the text underneath it. Recognition quality compounds.

The problem

recognition quality sounds simple until you look at real documents. The failure cases are where most of the difficulty — and most of the value — lives.

What we did

We treated it as a measurement problem first. Before changing the model, we built a labeled set of the hardest examples we could find, so we could tell whether any change actually helped.

From there, the work was iterative: adjust the model, score it on the hard set, keep what moved the number, and throw out what didn't. No change ships without evidence.

  • A benchmark built from real, difficult documents.
  • Targeted training on the failure cases.
  • Confidence signals so downstream systems know when to check.

Results

96.1%
on our hard set
−52%
error rate

Want to see it on your own documents? Open the playground or reach out — we're happy to run a sample with you.

OCRextractionaccuracy
[]TRY DOCEN

Run your hardest documentthrough Docen.

See the structured output for yourself, or reach the team at [email protected].