AI Document Redaction: Better model, harder test

Redaction Accuracy

One missed name in a due diligence review can become a GDPR breach. That risk is the reason AI document redaction exists, and it is why a redaction tool is judged first on what it misses. Our Smart Redaction tool, has found and redacted personal data across our client’s data rooms for years – names, addresses, account numbers, the PII that GDPR requires you to protect. What changed this year is how much of the model it catches.


The latest version beats the model it replaces in every one of the six languages we tested, and it catches more sensitive information than the redaction service built into Microsoft Azure.

Redaction accuracy: what changed, and by how much

The new model raises overall accuracy by 5.1 points over the previous version, and beats Microsoft Azure on every measure we tested.

Redaction Accuracy Chart

Three numbers describe how a redaction tool performs. Catch rate, or recall, is how much of the sensitive data in the document the tool finds. Precision is how rarely it redacts text that did not need it. The overall score, known as F1, combines the two. The new model finishes ahead of Azure on all three.


The numbers come from a tougher test than the one behind our earlier results. The old benchmark counted redaction as a success if the model caught even one character of an item. This one measures the share of characters it gets right, a harder and more realistic bar, because a half-redacted address still leaves personal data exposed. We did two things this cycle: improved the model through better training and held it to this stricter standard. That is why the figures read lower than the 93% we reported in 2025, while the jump from the previous model, measured the same hard way, is earned rather than a trick of measurement.

The number to watch: catch rate


For redaction, one number outranks the rest. The previous model caught 74% of the sensitive data in a document. The new one catches 84.6%.


That jump is the whole point. A tool can post a respectable overall score while missing one address in four, and in redaction, a miss is the failure that matters. Every character left unredacted is still sensitive information left on a page. Lifting the catch rate from 74% to nearly 85% removes a large share of those misses, and it pushes past Azure on the measure that carries the most compliance risk.


You might notice the new model’s precision sits a little below the old one, 85.4% against 87.2%. That is a deliberate trade, and the right one for redaction. When precision dips, the tool sometimes redacts a word that did not need it, and a reviewer can undo that in seconds. When catch rate dips, the miss slips through unnoticed until it is too late. Over-redaction is an annoyance. Under-redaction is a breach. We tuned the model to err on the safe side, and even so it redacts more precisely than Azure does.

The change was in the data, not the model

The underlying model is a multi-lingual large language model, the same one we used before. We did not reach for a bigger engine. We changed what we fed it.


AI models learn from examples. To redact well across many language and document types, a model needs to study a balanced set of examples from all of them. Real examples are the problem. Annotated legal documents full of personal data are sensitive by nature, slow to collect, and heavily skewed towards the language of a provider’s earliest customers. A model trained mostly on English or German contracts develops blind spots. Hand it a Danish or Finnish document and it stumbles, even though the personal data it needs to find is the same.

placeholder-wide

We fixed this with realistic synthetic data: documents we generated that look and read like the real thing. Once that generation pipeline exists, we control the balance. We decide how many documents come from each language and how often each time of information appears, then train on a set that covers all of them evenly.

Better in every language, and on the fields that are hardest

We tested the new model on 109 documents across six languages: German, English, French, Italian, Dutch, and Polish. It beats the previous model in every one.


The largest gains landed where redaction is hard. A phone number or an email address follows a fixed pattern a tool can spot on sight. Addresses, company names, and people’s names do not. Whether a word is a company name depends on the sentence around it. The model has to read the context and understand it, and that is where weaker tools miss. Our biggest improvements came on these three fields, which is where a redaction tool earns its place against competition.

Why we publish these numbers

Most virtual data room providers describe their redaction as accurate and leave it there. You are asked to blindly trust it. We take a different view. When a miss can mean a breach, you deserve to see how the tool performs before you rely on it. So, we publish our numbers, name the tool we benchmark against, and show the results. It is one of the questions worth putting to any provider, and we bring the same openness to the rest of our Smart VDR suite.

Frequently asked questions

Want to find out more?

Submit your details below and a member of our team will get back to you shortly

This field is for validation purposes and should be left unchanged.
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form