Docen
BLOG

Notes from the lab.

Releases, engineering write-ups, benchmarks, and stories from teams putting Docen to work.

Release·May 19, 2026

Docen Parse 2.1

Faster pages, better equations, and improved reading order on dense layouts.

Read
Release

Managed batch processing is now live

Point Docen at a queue of documents and let throughput scale itself.

Apr 8, 2026·5 min
Release

Docen OCR 2

Our recognition model gets better at faint scans, handwriting, and dense multilingual pages.

Mar 24, 2026·5 min
Release

Turbo extraction

A faster extraction path for high-volume, latency-sensitive workloads.

Feb 25, 2026·4 min
Product

EU data residency for Docen

Options for keeping document processing within your chosen region.

Feb 11, 2026·4 min
Benchmarks

Saturating the olmOCR benchmark

What it means when a benchmark stops being hard — and where we look next.

Jan 28, 2026·6 min
Product

Build document processing pipelines with workflows

Chain parsing, layout, and extraction into a repeatable pipeline.

Jan 14, 2026·6 min
Release

Introducing balanced extraction mode

A new mode that trades a little speed for higher recall on hard fields.

Dec 16, 2025·4 min
Product

Automatically fill PDF forms with AI

Read a form's fields, map your data, and produce a completed PDF.

Dec 3, 2025·5 min
Product

View your API requests in the playground

Inspect, replay, and share the exact requests you send to Docen.

Nov 25, 2025·3 min
Product

Extracting tracked changes and metadata

Pull revision history and hidden metadata out of documents, not just the visible text.

Nov 18, 2025·5 min
Benchmarks

How we benchmark and evaluate Docen

Our approach to measuring parsing and extraction quality — what we test, how we score it, and why we publish it.

Nov 5, 2025·7 min
Release

Docen Eval for extraction

Bring evaluation to your extraction schemas, field by field.

Oct 21, 2025·4 min
Launch week

Launch week: Docen Parse is faster again

Another round of speedups for the parsing model.

Oct 20, 2025·2 min
Launch week

Launch week: spreadsheet parsing

Parse spreadsheets with empty cells and irregular headers, reliably.

Oct 19, 2025·3 min
Launch week

Launch week: playground examples

A gallery of ready-made examples to explore what Docen can do.

Oct 18, 2025·2 min
Launch week

Launch week: a new playground

Try any processor on your own documents, right in the browser.

Oct 17, 2025·3 min
Launch week

Launch week: high-accuracy mode

A mode tuned for the documents where every character counts.

Oct 16, 2025·3 min
Launch week

Launch week: faster tracked-changes outputs

Tracked-changes extraction, now noticeably faster.

Oct 15, 2025·3 min
Launch week

Launch week: introducing Docen Layout for multi-page section hierarchy

Keeping section hierarchy intact across long, multi-page documents.

Oct 14, 2025·4 min
Launch week

Launch week: document segmentation

Split long documents into clean, addressable sections automatically.

Oct 13, 2025·3 min
Launch week

Launch week: the layout model

Region and table detection that holds together end to end.

Oct 12, 2025·3 min
Release

Introducing Docen Eval

Score parsing and extraction against your own labeled documents before you ship.

Oct 7, 2025·5 min
Release

The Docen SDKs

Official client libraries that drop Docen into the stack you already run.

Sep 30, 2025·4 min
Case study

How a presentation platform turned decks into structured data

A fast-growing presentation platform needed to read user-uploaded decks and documents reliably.

Sep 10, 2025·5 min
Product

Structured extraction with citations

Every extracted value comes with a pointer to the exact span it came from.

Sep 2, 2025·5 min
Product

Word bounding boxes and confidence

Every recognized word comes with a location and a confidence score. Here's how to use them.

Aug 27, 2025·4 min
Release

Introducing Docen Extract

Schema-driven structured extraction with citations back to the source span.

Aug 19, 2025·5 min
Engineering

Reducing hallucinations in document extraction

How citations and confidence keep extracted values grounded in the source.

Aug 5, 2025·7 min
Engineering

Speeding up Docen Parse

How we cut median page latency while improving accuracy.

Jul 29, 2025·6 min
Engineering

Structured extraction with the Docen API and long documents

Strategies for extracting from documents that don't fit in a single pass.

Jul 8, 2025·7 min
Case study

Turning exam papers into structured practice at scale

An exam-prep company used Docen to parse past papers, mark schemes, and diagrams into clean, structured content.

Jun 24, 2025·5 min
Release

Docen Parse 2

A new generation of the parsing model, with sharper structure and stronger handwriting.

Jun 17, 2025·6 min
Engineering

High-fidelity OCR drives accurate structured extraction

Extraction is only as good as the text underneath it. Recognition quality compounds.

Jun 3, 2025·6 min
Engineering

Cracking math OCR

Recognizing equations and notation is its own hard problem. Here's how we approach it.

May 27, 2025·7 min
Case study

Reading purchase orders and invoices without manual entry

A procurement team replaced manual data entry with Docen and cut invoice turnaround from days to minutes.

May 13, 2025·4 min
Engineering

Extracting hyperlinks from PDFs

Links are data too. Here's how Docen recovers them from PDFs.

Apr 15, 2025·4 min
Case study

Structured content for an AI learning company

An AI learning company needed clean, structured text from messy source material to power its tutoring models.

Apr 2, 2025·5 min
Product

Parse PDFs just the way you want

Shape the output format and structure to fit what your systems expect.

Mar 25, 2025·5 min
Case study

Segmenting benefits and policy documents for a healthcare team

A healthcare team used Docen to split dense benefits documents into addressable, searchable sections.

Mar 11, 2025·5 min
Case study

Digitizing a county historical archive, page by page

A county records office turned a century of handwritten archives into searchable, structured text.

Feb 18, 2025·6 min
Release

Docen Parse 1.5

A big step up in table accuracy and multilingual recognition.

Feb 4, 2025·4 min
Case study

Extracting part data from electronics datasheets

An electronics-sourcing platform used Docen to pull structured specs from thousands of component datasheets.

Jan 22, 2025·5 min
Company

Free to start, pay as you go

A simpler way to price document intelligence: start free, then pay for what you process.

Jan 9, 2025·3 min
Launch week

Launch week: Docen Parse 1.1

Better tables, faster pages, and steadier reading order across long documents.

Nov 12, 2024·3 min
Release

Introducing Docen Parse

Our core parsing model turns any document into clean, ordered Markdown, HTML, and JSON.

Sep 24, 2024·5 min
Company

Building better document intelligence for an AI-first world

Why accurate parsing and extraction are becoming core infrastructure — and how we're building models for the documents that matter.

Jul 16, 2024·6 min