parsing enginedocuments in, data out

Any document.
Clean structured data.

Datahone parses invoices, PDFs, CVs, spreadsheets, emails, and scanned images into structured JSON your tools can use — reading layout, context, and meaning with the latest vision language models, not just the characters on the page.

Any file formatStructured JSON out60+ languages
Parsed documents become structured data — routed where your team already works. Connect via webhook or the REST API.
what is document parsing

Parsing turns a document into data — not just text.

Document parsing is the process of reading a file and pulling out the information that matters as structured, labelled data — the vendor, the total, the date, the line items — ready to drop straight into your systems. It's the difference between a page you have to read and a record you can act on.

Datahone reads any document — typed or handwritten, clean or scanned — and interprets its layout, context, and meaning with the latest vision language models. Optical character recognition is one step inside that pipeline; understanding the document is the rest. What comes back is clean JSON, tables included, routed to the tools you already use.

60+languages read

Latin and non-Latin scripts, with experimental support well beyond.

Any formatone engine

PDFs, images, emails, and spreadsheets — typed or handwritten — parsed in a single pipeline.

Confidence scoringon every field

Each value comes back with a confidence score and a built-in review step, so your team can verify at a glance.

every document type

One engine parses every kind of document.

Invoices, PDFs, CVs, spreadsheets, emails, scanned images — whatever lands, datahone reads it, understands it, and returns it as structured data.

01

Invoices & receipts

Totals, tax, dates, supplier details, and every line item pulled from any invoice or receipt layout. See the invoice page.

02

PDFs, searchable & scanned

Text-layer PDFs are read from source; image-only scans go through computer-vision OCR. Either way you get structure. See the PDF page.

03

CVs & résumés

Contact details, work history, skills, and education parsed into consistent records for your ATS. See the CV page.

04

Emails & attachments

Parse email bodies — including rich HTML — and the documents attached to them in one pass. See the email page.

05

Spreadsheets & tables

XLSX and CSV files parsed as rows — clean records, not flattened text — with tables preserved. See the spreadsheet page.

06

Photos & scanned images

A phone photo, a JPG, a TIFF, a fax — image-only files are read with OCR and handwriting recognition, then parsed like any other document.

printed + handwritten
how it works

From any document to structured data in three steps.

Send in a file, let datahone read and understand it, and receive clean data where you need it. You only ever touch three surfaces — the rest is automatic.

Step 01 · ingest

Send in any document

Upload, forward, or POST a PDF, image, scan, email, or spreadsheet. Datahone detects the format and routes image-only pages through computer-vision OCR automatically.

invoice_1224.pdfscanned image · reading…
Step 02 · parse

AI reads and understands it

The latest vision language models read the document the way a person would — telling fields apart, following tables, handling layouts that shift. Each value comes back with a confidence score.

vendorAyemarine Ltd
total$163.00
date2024-03-14
confidence0.98
Step 03 · structure

Get structure, not a text dump

Fields and tables come back as clean JSON — POSTed to any webhook you run, pulled from the REST API, or downloaded as CSV or XLSX — with a confidence score on every value, so your team can verify at a glance.

POST/datahone/webhook200 ok
GETapi/v1/documents/:id/output200 ok
GETapi/v1/exports?format=csv202 queued
no setup, no rules

You describe the data. Datahone builds the template.

There are no zones to draw or rules to maintain. You simply chat with datahone about the fields you need, it confirms them against your document, and a reusable template for that document type is created for you automatically — powered by the latest vision language models reading layout, context, and meaning.

Describe it in a chat

Tell datahone the fields you want in plain language. It reads a sample, confirms what it found, and you agree the schema in a short conversation — no forms, no field mapping.

conversational setup

Templates build themselves

Once you've confirmed the fields, datahone generates a reusable template for that document type automatically and matches every future document to it — nothing to hand-build or maintain.

auto-generated

Layouts that shift, handled

Because templates come from understanding rather than fixed coordinates, fields that move, resize, or repeat between documents — variable line-item tables, mixed layouts — are followed automatically.

no fixed zones
Every template is created for you from a short conversation — there's nothing to draw, map, or maintain. Need a specific document type? See the invoice, PDF, and spreadsheet pages.
questions

Document parsing, answered.

Q · 01

Document parsing reads a file and pulls out the information that matters as structured, labelled data — vendor, total, date, line items — rather than a flat block of text. Datahone parses invoices, PDFs, CVs, spreadsheets, emails, and scanned images into clean JSON your systems can use.

Q · 02

PDFs (searchable and scanned), images (JPG, PNG, TIFF, photos), emails and their attachments, DOCX documents, and spreadsheets (XLSX, CSV). Legacy binary .doc and .xls files are not parsed — they are declined up front with a clear unsupported-format response. Image-only files are read with OCR and handwriting recognition first, then parsed like any other document.

Q · 03

OCR converts pixels into text. Parsing goes further — it works out what that text means: which string is the date, which is the total, which rows belong to a table. Datahone runs OCR where it's needed and the latest vision language models on top, so you get labelled data, not a transcript. See the OCR software page.

Q · 04

Clean, structured JSON — fields and tables, with a confidence score on every value. It's delivered to any webhook you run, pulled from the REST API, or downloaded as JSON, CSV, or XLSX, ready for your team to route into the tools you already use.

Q · 05

More than 60 languages, including English, Spanish, French, German, Dutch, Russian, Japanese, Korean, Chinese, Hebrew, Arabic, and Hindi, with experimental support for another 160+. Both Latin and non-Latin scripts are handled.

Q · 06

No. You describe the fields you want in a short chat, datahone confirms them against a sample document, and it creates a reusable template for that document type automatically. There are no zones to draw, no fields to map, and nothing to maintain as your documents change.

Q · 07

Every value comes back with a confidence score, so you always know how reliable a read is. High-confidence values flow straight through, and a built-in review step lets your team confirm anything you want a second pair of eyes on before it reaches your systems.

Q · 08

Per page, with a hard cap on every tier. PDFs, scans, and images count by their page count; spreadsheets count one page per 10,000 cells. When you hit your limit the next document is blocked — never silently billed. See pricing →

get started

Parse your first document free.

Drop in an invoice, a PDF, a CV, or a spreadsheet and watch datahone hand it back as structured data — fields and tables, not a text dump. Free in minutes, no card required.

Any document typeStructured JSON outNo card required