At a glance:
- We benchmarked the July 2026 release of Imprima’s Smart Redaction against Azure Redaction (used by most data rooms) and GPT-5.6, the latest general-purpose GPT model, on the same 80 documents.
- Smart Redaction is more accurate than Azure: it catches 78% of sensitive data to Azure’s 62%, so Azure leaves nearly twice as much personal data exposed. Its overall score is 0.84 against 0.66.
- Smart Redaction beats Azure in every language tested, including English and German, and in lower-resource languages such as Polish and Czech.
- It also beats GPT-5.6 on accuracy, and runs about five times faster per document.
- On the hardest fields, where redaction depends on context e.g. people’s names, company names and addresses, Smart Redaction beats both Azure and GPT-5.6.
We keep improving the AI behind your data room
The AI in an Imprima data room is not a fixed feature that we shipped once and left alone. We improve our models continuously, and Smart Redaction, the tool that finds and protects personal data before a document is shared, is no exception.
With the latest release, we wanted to know exactly where it stands, so we measured it.
What we benchmarked, and against which models
We took the July 2026 release of Smart Redaction and tested it against two powerful tools relevant for this market:
- Azure AI language redaction, the engine most other virtual data room providers rely on for their redaction.
- GPT-5.6, the latest general-purpose GPT model at the time of testing1.
Every model saw the same 80 documents2 and was scored the same way: character by character, measuring the share of sensitive characters each one redacted. That is the strict, real-world test, because a half-redacted address still leaves personal data on the page.
How to read the numbers
Three numbers describe how a redaction tool performs.
- Catch rate, or recall, is how much of the sensitive data in the document the tool finds.
- Precision is how rarely it redacts text that did not need it.
- The overall score, known as F1, combines the two.
Recall is a true proportion, so we give it as a percentage in the text: a recall of 78% means the tool caught 78% of the sensitive characters. F1 is a combined score rather than a share of anything, so we give it on a 0-to-1 scale. The charts plot on the same 0-to-1 axis, where a recall of 78% appears as 0.78.
For redaction, recall is the one to watch because whatever the tool misses stays exposed, which is the exact thing redaction aims to prevent. Precision matters too: over-redacting makes a document harder to read, but a stray black box is a minor inconvenience a reviewer can undo in seconds, whereas a missed detail is a leak no one may notice. That is why the charts below lead on recall and report F1 next to it, as a check that catching more has not come at too high a price in precision.
Smart Redaction performs better than Azure
Measured over the whole set, Smart Redaction leads Azure on every metric.
- Smart Redaction
- Microsoft Azure
At a recall rate of 0.62, Azure misses around 38% of the sensitive characters, compared to Smart Redaction which misses around 22%. In other words, Azure leaves nearly twice as much personal data on the page.
English and German
Broken down by language, the pattern holds. In English and German, Smart Redaction remains well ahead of Azure.
- Smart Redaction
- Microsoft Azure
In German, the F1 gap is 0.91 against 0.71, and in English 0.84 against 0.67, a very decisive margin in both.
Polish and Czech
The advantage holds in Polish and Czech too, languages for which far less training data is generally available.
- Smart Redaction
- Microsoft Azure
In Polish, Smart Redaction reaches an F1 of 0.82 against Azure’s 0.72. Czech is a single document in this run, so reads as indicative rather than conclusive, but it points the same way: 0.91 against 0.84.
Smart Redaction even performs better than GPT-5.6, and is much faster
A general-purpose LLM prompted to find personal data is the other obvious tool to reach for, so we tested the latest one on the same 80 documents.
- Smart Redaction
- GPT-5.6 (Terra)
GPT-5.6 is a much stronger redactor than Azure. But on the measure that decides a redaction tool, recall, it falls behind: 0.70 against 0.78, so it leaves noticeably more sensitive data exposed. Its overall F1 trails too, 0.78 to 0.84.
Accuracy aside, a general LLM is slow for redaction at data room scale. In our timing run, GPT-5.6 averaged about five times slower per document than Smart Redaction. Across a data room of thousands of pages, that’s the difference between a batch you can share quickly, and one that slows the deal down.
The hardest fields to redact
Speed aside, the accuracy picture has one more layer worth opening up. Some kinds of personal data follow a fixed pattern that a tool can match on sight, such as an email address or phone number, and all three tools handle those well. Others are harder. To redact a person’s name, a company name or an address, the tool has to read the surrounding text and work out where the entity starts and ends, because the same word can be a name in one sentence and something ordinary in the next.
These context dependent fields are where redaction is won or lost, and they are where Smart Redaction leads both Azure and GPT-5.6 (F1 scores below).
- Smart Redaction
- Microsoft Azure
- GPT-5.6 (Terra)
On person names, Smart Redaction scores 0.93 ahead of GPT-5.6 (0.91) and Azure (0.87). On addresses it reaches 0.88 against 0.78 and 0.81. The widest gap is on company names where Smart Redaction scores 0.81 while GPT-5.6 and Azure sit at 0.67 and 0.70. These are the fields most likely to leak in a real document, and the ones a model trained for the task handles better than either Azure or GPT-5.6
Why we publish these results
Most data room providers describe their redaction as accurate and leave it there. We would rather show you the numbers, name the tools we benchmark against, and let you see how we compare. We bring the same openness to the rest of our Smart VDR suite too.
See how Smart Redaction performs on your own documents – Book a Demo
Footnotes