ocr enginescan in, structure out

Accurate OCR.
Structured output.

Datahone's OCR engine reads printed and handwritten text from PDFs, scans, images, emails, and spreadsheets — in 60+ languages — then returns clean, structured fields and tables, ready for your tools.

Printed & handwrittenScanned PDFs & images60+ languages
Recognised text becomes structured data — routed where your team already works. Connect via webhook or the REST API.
what is ocr software

OCR reads the page. Datahone understands it.

Optical character recognition is the technology that converts the characters in a scan, photo, or PDF into machine-readable text — the foundation every document workflow is built on. Most OCR tools stop there and hand you the raw text.

Datahone goes further. Our engine pairs high-accuracy recognition with the latest vision language models, interpreting layout, context, and meaning the way a person would — telling dates from totals, names from line items, headers from handwriting. The result isn't a transcript you still have to read; it's clean, structured data your systems can act on the moment it lands.

60+languages read

Latin and non-Latin scripts, with experimental support well beyond.

Printed & handwrittenboth supported

Typed pages and handwriting are read in the same pass — no upcharge.

Confidence scoringon every field

Each value comes back with a confidence score and a built-in review step, so your team can verify at a glance.

ocr for all

One engine reads every kind of document.

Searchable PDFs, image-only scans, handwriting, emails, spreadsheets — whatever lands, datahone recognises the text and returns it as structure.

01

Text-based PDFs

When a PDF already carries a text layer — searchable PDFs and PDF/A — datahone reads it straight from the source for fast, faithful results.

02

Scanned PDFs & images

For pages that are just images with no text layer, computer-vision OCR recognises and extracts the text — JPG, PNG, TIFF, and scanned PDFs alike.

03

Handwriting recognition

Read handwritten text in Latin, Japanese, and Korean alphabets, with experimental support for more — recognised in the same pass as printed pages.

printed + handwritten
04

Emails & text documents

Recognise text in email bodies — including rich HTML emails with images and links — plus DOCX documents and other text files in one pass.

05

Spreadsheets & more

Read XLSX and CSV files, web pages, and other formats — and pull tables back as structured rows, not flattened text. See how a page is counted.

06

60+ languages

Trained to recognise text in more than 60 languages — English, Spanish, French, German, Russian, Japanese, Korean, Chinese, Arabic, Hindi and more — with experimental support for 160+ beyond.

how it works

From a scan to structured data in three steps.

OCR is the first stage, not the finished result. Datahone reads the characters, applies the latest vision language models where a page needs them, locates the fields, and delivers clean structured data — you only ever touch three surfaces.

Step 01 · scan

Send in the document

Upload, forward, or POST a PDF, image, scan, email, or spreadsheet. Datahone detects whether a text layer exists and routes image-only pages through computer-vision OCR.

invoice_1224.pdfscanned image · reading…
Step 02 · recognise

OCR reads every character

Printed or handwritten, across 60+ languages, the engine recognises the text and locates it on the page — drawing on the latest vision language models where a page calls for deeper interpretation. Each value comes back with a confidence score.

vendorAyemarine Ltd
total$163.00
date2024-03-14
confidence0.98
Step 03 · structure

Get structure, not a text dump

Fields and tables come back as clean JSON — POSTed to any webhook you run, pulled from the REST API, or downloaded as CSV or XLSX — with a confidence score on every value, so your team can verify at a glance.

POST/datahone/webhook200 ok
GETapi/v1/documents/:id/output200 ok
GETapi/v1/exports?format=csv202 queued
no setup, no rules

You describe the data. Datahone builds the template.

There are no zones to draw or rules to maintain. You simply chat with datahone about the fields you need, it confirms them against your document, and a reusable template for that document type is created for you automatically — powered by the latest vision language models reading layout, context, and meaning.

Describe it in a chat

Tell datahone the fields you want in plain language. It reads a sample, confirms what it found, and you agree the schema in a short conversation — no forms, no field mapping.

conversational setup

Templates build themselves

Once you’ve confirmed the fields, datahone generates a reusable template for that document type automatically and matches every future document to it — nothing to hand-build or maintain.

auto-generated

Layouts that shift, handled

Because templates come from understanding rather than fixed coordinates, fields that move, resize, or repeat between documents — variable line-item tables, mixed layouts — are followed automatically.

no fixed zones
Every template is created for you from a short conversation — there's nothing to draw, map, or maintain. Need a specific document type? See the invoice, PDF, and spreadsheet pages.
questions

OCR software, answered.

Q · 01

OCR — optical character recognition — reads the characters in a scan, photo, or PDF and converts them into machine-readable text. Datahone runs OCR as the first step of a pipeline that also finds the structure, so you get usable fields and tables rather than a raw text dump.

Q · 02

Yes. When a PDF has no text layer, or you upload a JPG, PNG, TIFF, or photo of a page, computer-vision OCR recognises and extracts the text. When a text layer is already present, datahone reads it straight from the source.

Q · 03

Handwritten text is supported in Latin, Japanese, and Korean alphabets, with experimental support for others, and is recognised in the same pass as printed pages. Each entry comes back with a confidence score and a built-in review step, so your team can confirm anything worth a second look.

Q · 04

More than 60 languages, including English, Spanish, French, German, Dutch, Russian, Japanese, Korean, Chinese, Hebrew, Arabic, and Hindi, with experimental support for another 160+. Both Latin and non-Latin scripts are handled.

Q · 05

Both. The OCR step recognises the text; the extraction step turns it into structured fields and tables — dates, totals, names, line items — as clean JSON. You define what you need in a short conversation and datahone builds the template for you, so getting usable data takes no separate tool and no manual setup.

Q · 06

No. You describe the fields you want in a short chat, datahone confirms them against a sample document, and it creates a reusable template for that document type automatically. There are no zones to draw, no fields to map, and nothing to maintain as your documents change.

Q · 07

Every value comes back with a confidence score, so you always know how reliable a read is. High-confidence values flow straight through, and a built-in review step lets your team confirm anything you want a second pair of eyes on before it reaches your systems.

Q · 08

Per page, with a hard cap on every tier. PDFs, scans, and images count by their page count; CSV and XLSX files count one page per 10,000 cells. When you hit your limit the next document is blocked — never silently billed. See pricing →

get started

Run an OCR you can build on.

Drop in a scan, an image, or a PDF and watch datahone hand it back as structured data — fields and tables, not a text dump. Free in minutes, no card required.

Printed & handwritten60+ languagesStructured output