Detect the PII
Send a file — or chain it straight after extraction. datahone scans text, tables, and documents and flags every field that identifies a person, including mentions buried in free text.
Run documents, tables, and logs through datahone's PII pass and personal data is detected, then redacted, masked, or pseudonymised — so your team can share datasets with vendors, analysts, and AI tools without leaking anyone's identity. EU-hosted, never used to train shared models.
Most data your team touches carries personal details — names, emails, phone numbers, addresses, ID and card numbers, dates of birth, and stray mentions buried in free text. The moment you share it with a vendor, an analyst, or an AI tool, that personal data goes with it. Anonymisation removes the link to a real person while keeping the data useful.
datahone detects PII across text, tables, and documents, then applies a policy you choose per field — redact it, mask it, or pseudonymise it with a consistent token that stays stable across files so your joins still work. It runs as a step after extraction, with an audit trail of everything changed.
Point datahone at the data you need to share, set a policy once, and every file after that comes back safe — with a record of exactly what changed. You only ever touch three surfaces.
Send a file — or chain it straight after extraction. datahone scans text, tables, and documents and flags every field that identifies a person, including mentions buried in free text.
Choose what happens to each field — redact, mask, or pseudonymise with a consistent token. Low-confidence detections wait in a review queue rather than being passed through.
The anonymised dataset — plus a full audit trail — is POSTed to any webhook you run, pulled from the REST API, or downloaded as JSON, CSV, or XLSX — ready for you to route into a vendor handoff, an analytics set, or your warehouse. Review it in the dashboard first.
A scan line crosses a real record. As it passes, every field that identifies a person is rewritten — redacted, masked, or pseudonymised by policy — and the clean, safe-to-share payload assembles on the right. Country, company, and amounts are kept; the identity is not.
From a spreadsheet of customers to a folder of contracts to a stream of support logs, datahone finds the personal data and applies the policy you set — consistently, with a trail of what changed.
Names, emails, phone numbers, addresses, national and ID numbers, card numbers, dates of birth, and free-text mentions — detected across documents, tables, and logs.
Replace a value with [REDACTED], mask all but the last digits, or swap it for a stable token. Pick the transform that keeps the data useful for the job.
The same value always maps to the same pseudonym, across every file you run — so de-identified datasets still join on a person without ever exposing who they are.
Set a different rule for every field, save it as a reusable policy, and apply it to every file of that shape — so the same dataset is anonymised the same way each time.
Every change is logged — field, policy, and time. An optional encrypted vault keeps the mapping, so authorised users can reverse a pseudonym when there is a lawful reason to.
PDFs, DOCX files, scans, CSV and XLSX tables, emails, and JSON logs all go through the same PII pass — and come back in the same shape, minus the personal data.
It turns 'we can't share that, it has personal data in it' into a step that runs itself — so the data moves, and the identities don't.
Built to support GDPR, the UK Data Protection Act, and CCPA. Minimise personal data before it leaves your systems, with an audit trail to show what was changed.
Your files and the data inside them are encrypted in transit and at rest, stored on EU-hosted infrastructure, and scoped to your account alone.
Your content is processed to anonymise it and nothing more. It is never used to train shared models, and you can set retention so it isn't kept longer than you need.
We never claim to catch 100% of PII. Fields the model isn't sure about are held in a review queue for a person to confirm — never silently passed through.
Anonymisation runs on data you've already pulled out of a document. The same engine does the extraction — point it at the source first, then scrub. See everything datahone extracts →
Pull fields, tables, and totals out of any PDF — digital, scanned, or handwritten — then send the result through the PII pass.
Explore SolutionSupplier, line items, totals, VAT, dates, and PO numbers — straight from any invoice, ready to anonymise before sharing.
Explore SolutionTurn a messy CSV or XLSX workbook into clean, typed, de-duplicated rows — then de-identify the columns that carry PII.
ExploreAnything that identifies a person: names, emails, phone numbers, postal addresses, national and other ID numbers, payment card numbers, dates of birth — and mentions of any of these inside free-text fields and document bodies, not just neatly labelled columns.
Redact replaces the value entirely with [REDACTED]. Mask keeps a usable shape but hides the detail — like +44 7•• ••• •12. Pseudonymise swaps it for a consistent token such as usr_8f3a, so the data still joins across records without revealing the person.
Yes. Each field gets its own rule — redact the name, mask the phone, pseudonymise the email, keep the country. Save that as a reusable policy and apply it to every file of the same shape so anonymisation is consistent across runs.
They do. The same input value always maps to the same token within your account, so a person who appears in two different files gets the same pseudonym in both — letting de-identified datasets still be joined without exposing identities.
Redaction and masking are not — the original is gone from the output. Pseudonymisation can be reversible if you enable the optional encrypted vault, which stores the token-to-value mapping so authorised users can re-identify when there's a lawful basis. Leave the vault off and the mapping is never kept.
We don't claim to catch every instance. Detections the model isn't confident about are flagged and held in a review queue for a person to confirm or correct, rather than guessed and written through — so uncertain fields fail safe instead of leaking.
After. datahone first extracts structured data from your documents or tables, then the PII pass runs on the result as a separate step. You can also send already-extracted data straight to anonymisation through the API. Either way it's EU-hosted and never used to train shared models.
Per page, the same as the rest of datahone — documents by page count, tables at one page per 10,000 cells. Every tier has a hard cap, so a large dataset is blocked rather than silently over-billed. See pricing →
Drop in a file and watch datahone hand it back with the personal data redacted, masked, or pseudonymised to your policy — ready to share, with a full audit trail. Free in minutes, no card required.