How we benchmark and evaluate Docen
Our approach to measuring parsing and extraction quality — what we test, how we score it, and why we publish it.
Our approach to measuring parsing and extraction quality — what we test, how we score it, and why we publish it.
What we measure
We publish benchmarks because a document intelligence model is only as good as its behavior on documents you can't cherry-pick. We test recognition, tables, layout, and extraction against sensible baselines.
- Character and word error rate on scanned and photographed pages.
- Table-cell accuracy on complex, merged layouts.
- Field-level F1 for schema extraction, with citations checked.
- Latency per page at production settings.
Results
Method
Every number here comes from documents held out of training. We report the settings, keep the evaluation reproducible, and update the figures as models change. When a result looks too good, we assume the test is wrong until we've checked it.
Want to see it on your own documents? Open the playground or reach out — we're happy to run a sample with you.