Text-based PDFs
When a PDF already carries a text layer — searchable PDFs and PDF/A — datahone reads it straight from the source for fast, faithful results.
Datahone's OCR engine reads printed and handwritten text from PDFs, scans, images, emails, and spreadsheets — in 60+ languages — then returns clean, structured fields and tables, ready for your tools.
Optical character recognition is the technology that converts the characters in a scan, photo, or PDF into machine-readable text — the foundation every document workflow is built on. Most OCR tools stop there and hand you the raw text.
Datahone goes further. Our engine pairs high-accuracy recognition with the latest vision language models, interpreting layout, context, and meaning the way a person would — telling dates from totals, names from line items, headers from handwriting. The result isn't a transcript you still have to read; it's clean, structured data your systems can act on the moment it lands.
Latin and non-Latin scripts, with experimental support well beyond.
Typed pages and handwriting are read in the same pass — no upcharge.
Each value comes back with a confidence score and a built-in review step, so your team can verify at a glance.
Searchable PDFs, image-only scans, handwriting, emails, spreadsheets — whatever lands, datahone recognises the text and returns it as structure.
When a PDF already carries a text layer — searchable PDFs and PDF/A — datahone reads it straight from the source for fast, faithful results.
For pages that are just images with no text layer, computer-vision OCR recognises and extracts the text — JPG, PNG, TIFF, and scanned PDFs alike.
Read handwritten text in Latin, Japanese, and Korean alphabets, with experimental support for more — recognised in the same pass as printed pages.
printed + handwrittenRecognise text in email bodies — including rich HTML emails with images and links — plus DOCX documents and other text files in one pass.
Read XLSX and CSV files, web pages, and other formats — and pull tables back as structured rows, not flattened text. See how a page is counted.
Trained to recognise text in more than 60 languages — English, Spanish, French, German, Russian, Japanese, Korean, Chinese, Arabic, Hindi and more — with experimental support for 160+ beyond.
OCR is the first stage, not the finished result. Datahone reads the characters, applies the latest vision language models where a page needs them, locates the fields, and delivers clean structured data — you only ever touch three surfaces.
Upload, forward, or POST a PDF, image, scan, email, or spreadsheet. Datahone detects whether a text layer exists and routes image-only pages through computer-vision OCR.
Printed or handwritten, across 60+ languages, the engine recognises the text and locates it on the page — drawing on the latest vision language models where a page calls for deeper interpretation. Each value comes back with a confidence score.
Fields and tables come back as clean JSON — POSTed to any webhook you run, pulled from the REST API, or downloaded as CSV or XLSX — with a confidence score on every value, so your team can verify at a glance.
There are no zones to draw or rules to maintain. You simply chat with datahone about the fields you need, it confirms them against your document, and a reusable template for that document type is created for you automatically — powered by the latest vision language models reading layout, context, and meaning.
Tell datahone the fields you want in plain language. It reads a sample, confirms what it found, and you agree the schema in a short conversation — no forms, no field mapping.
conversational setupOnce you’ve confirmed the fields, datahone generates a reusable template for that document type automatically and matches every future document to it — nothing to hand-build or maintain.
auto-generatedBecause templates come from understanding rather than fixed coordinates, fields that move, resize, or repeat between documents — variable line-item tables, mixed layouts — are followed automatically.
no fixed zonesOCR — optical character recognition — reads the characters in a scan, photo, or PDF and converts them into machine-readable text. Datahone runs OCR as the first step of a pipeline that also finds the structure, so you get usable fields and tables rather than a raw text dump.
Yes. When a PDF has no text layer, or you upload a JPG, PNG, TIFF, or photo of a page, computer-vision OCR recognises and extracts the text. When a text layer is already present, datahone reads it straight from the source.
Handwritten text is supported in Latin, Japanese, and Korean alphabets, with experimental support for others, and is recognised in the same pass as printed pages. Each entry comes back with a confidence score and a built-in review step, so your team can confirm anything worth a second look.
More than 60 languages, including English, Spanish, French, German, Dutch, Russian, Japanese, Korean, Chinese, Hebrew, Arabic, and Hindi, with experimental support for another 160+. Both Latin and non-Latin scripts are handled.
Both. The OCR step recognises the text; the extraction step turns it into structured fields and tables — dates, totals, names, line items — as clean JSON. You define what you need in a short conversation and datahone builds the template for you, so getting usable data takes no separate tool and no manual setup.
No. You describe the fields you want in a short chat, datahone confirms them against a sample document, and it creates a reusable template for that document type automatically. There are no zones to draw, no fields to map, and nothing to maintain as your documents change.
Every value comes back with a confidence score, so you always know how reliable a read is. High-confidence values flow straight through, and a built-in review step lets your team confirm anything you want a second pair of eyes on before it reaches your systems.
Per page, with a hard cap on every tier. PDFs, scans, and images count by their page count; CSV and XLSX files count one page per 10,000 cells. When you hit your limit the next document is blocked — never silently billed. See pricing →
Drop in a scan, an image, or a PDF and watch datahone hand it back as structured data — fields and tables, not a text dump. Free in minutes, no card required.