An applied document intelligence lab.
Docen builds models and tooling for the unglamorous, high-stakes work of turning real documents into accurate, structured data.
Most important information still lives in documents built for people, not machines: scanned records, dense filings, forms, spreadsheets, and PDFs that never quite parse cleanly. The teams that depend on that information spend enormous effort wrestling it into a usable shape.
We started Docen to close that gap with models trained specifically for document work — recognition, layout, parsing, and extraction — rather than general-purpose tools bent to fit. The goal is simple: output you can trust enough to build on.
We're a focused research and engineering group. We ship models, measure them in the open, and work closely with the teams that run them in production.
How we work.
Accuracy over hype
The number that matters is how often the output is right on the documents you actually process. We publish our benchmarks and hold ourselves to them.
Structure, not just text
A wall of characters isn't an answer. We preserve reading order, tables, and hierarchy so the output is usable downstream.
Run it anywhere
Managed cloud, your VPC, or fully offline — the same models and the same outputs, so where your data lives is your decision, not ours.
Measure everything
Every model ships with an evaluation story. If we can't measure a gain, we don't claim it.
The teams Docen is built for.
We work with teams across regulated, document-heavy fields — the places where a wrong field or a dropped table has real consequences.
Workloads range from a single hard scan to multi-million-page archives. If your documents are difficult, they're exactly the kind we build for. Browse use cases.
See what accurate looks likeon your documents.
Send us a sample of the documents you work with and we'll show you the structured output.