AI Document Redaction: How Smart Redaction compares to Azure and GPT-5.6

Smart Redaction Comparison

At a glance:

  • We benchmarked the July 2026 release of Imprima’s Smart Redaction against Azure Redaction (used by most data rooms) and GPT-5.6, the latest general-purpose GPT model, on the same 80 documents.
  • Smart Redaction is more accurate than Azure: it catches 78% of sensitive data to Azure’s 62%, so Azure leaves nearly twice as much personal data exposed. Its overall score is 0.84 against 0.66.
  • Smart Redaction beats Azure in every language tested, including English and German, and in lower-resource languages such as Polish and Czech.
  • It also beats GPT-5.6 on accuracy, and runs about five times faster per document.
  • On the hardest fields, where redaction depends on context e.g. people’s names, company names and addresses, Smart Redaction beats both Azure and GPT-5.6.

We keep improving the AI behind your data room

The AI in an Imprima data room is not a fixed feature that we shipped once and left alone. We improve our models continuously, and Smart Redaction, the tool that finds and protects personal data before a document is shared, is no exception.

With the latest release, we wanted to know exactly where it stands, so we measured it.

What we benchmarked, and against which models

We took the July 2026 release of Smart Redaction and tested it against two powerful tools relevant for this market:

  • Azure AI language redaction, the engine most other virtual data room providers rely on for their redaction.
  • GPT-5.6, the latest general-purpose GPT model at the time of testing1.

Every model saw the same 80 documents2 and was scored the same way: character by character, measuring the share of sensitive characters each one redacted. That is the strict, real-world test, because a half-redacted address still leaves personal data on the page.

How to read the numbers

Three numbers describe how a redaction tool performs.

  1. Catch rate, or recall, is how much of the sensitive data in the document the tool finds.
  2. Precision is how rarely it redacts text that did not need it.
  3. The overall score, known as F1, combines the two.

Recall is a true proportion, so we give it as a percentage in the text: a recall of 78% means the tool caught 78% of the sensitive characters. F1 is a combined score rather than a share of anything, so we give it on a 0-to-1 scale. The charts plot on the same 0-to-1 axis, where a recall of 78% appears as 0.78.

For redaction, recall is the one to watch because whatever the tool misses stays exposed, which is the exact thing redaction aims to prevent. Precision matters too: over-redacting makes a document harder to read, but a stray black box is a minor inconvenience a reviewer can undo in seconds, whereas a missed detail is a leak no one may notice. That is why the charts below lead on recall and report F1 next to it, as a check that catching more has not come at too high a price in precision.

Smart Redaction performs better than Azure

Measured over the whole set, Smart Redaction leads Azure on every metric.

  • Smart Redaction
  • Microsoft Azure
Catch rate (recall)
0.78
0.62
Overall (F1)
0.84
0.66
00.250.500.751.00
Scores across the full test set, all 80 documents, every language and field combined. Catch rate is the share of sensitive data a tool finds, so Smart Redaction at 0.78 against Azure’s 0.62 leaves close to half as much personal data exposed. The higher F1 (0.84 vs 0.66) shows that lead holds once precision is taken into account, so it is not just redacting more aggressively at the expense of precision.

At a recall rate of 0.62, Azure misses around 38% of the sensitive characters, compared to Smart Redaction which misses around 22%. In other words, Azure leaves nearly twice as much personal data on the page.

English and German

Broken down by language, the pattern holds. In English and German, Smart Redaction remains well ahead of Azure.

  • Smart Redaction
  • Microsoft Azure
English
Catch rate (recall)
0.79
0.70
Overall (F1)
0.84
0.67
German
Catch rate (recall)
0.87
0.60
Overall (F1)
0.91
0.71
00.250.500.751.00
The overall lead is not an average hiding a weak language. Broken out by language, Smart Redaction is ahead of Azure in both English and German, so a deal team working in either one sees the same advantage, not just those working across the full mix.

In German, the F1 gap is 0.91 against 0.71, and in English 0.84 against 0.67, a very decisive margin in both.

Polish and Czech

The advantage holds in Polish and Czech too, languages for which far less training data is generally available.

  • Smart Redaction
  • Microsoft Azure
Polish
Catch rate (recall)
0.74
0.64
Overall (F1)
0.82
0.72
Czech (1 document)
Catch rate (recall)
0.90
0.82
Overall (F1)
0.91
0.84
00.250.500.751.00
Lower-resource languages are where general tools usually fall off, because there is far less data to learn from. Smart Redaction holds its lead over Azure here too, so the advantage is not confined to the most common languages. (Czech is a single document, so read it as indicative.)

In Polish, Smart Redaction reaches an F1 of 0.82 against Azure’s 0.72. Czech is a single document in this run, so reads as indicative rather than conclusive, but it points the same way: 0.91 against 0.84.

Smart Redaction even performs better than GPT-5.6, and is much faster

A general-purpose LLM prompted to find personal data is the other obvious tool to reach for, so we tested the latest one on the same 80 documents.

  • Smart Redaction
  • GPT-5.6 (Terra)
Catch rate (recall)
0.78
0.70
Overall (F1)
0.84
0.78
00.250.500.751.00
We benchmarked against GPT-5.6 as a yardstick for what today’s most capable general-purpose AI can do. Smart Redaction still catches more of the sensitive data (0.78 vs 0.70), leaving less exposed, and comes out ahead on the overall F1 score too (0.84 vs 0.78), which balances how much it catches against how little it over-redacts. A model built for redaction does the job better.

GPT-5.6 is a much stronger redactor than Azure. But on the measure that decides a redaction tool, recall, it falls behind: 0.70 against 0.78, so it leaves noticeably more sensitive data exposed. Its overall F1 trails too, 0.78 to 0.84.

Accuracy aside, a general LLM is slow for redaction at data room scale. In our timing run, GPT-5.6 averaged about five times slower per document than Smart Redaction. Across a data room of thousands of pages, that’s the difference between a batch you can share quickly, and one that slows the deal down.

The hardest fields to redact

Speed aside, the accuracy picture has one more layer worth opening up. Some kinds of personal data follow a fixed pattern that a tool can match on sight, such as an email address or phone number, and all three tools handle those well. Others are harder. To redact a person’s name, a company name or an address, the tool has to read the surrounding text and work out where the entity starts and ends, because the same word can be a name in one sentence and something ordinary in the next.

These context dependent fields are where redaction is won or lost, and they are where Smart Redaction leads both Azure and GPT-5.6 (F1 scores below).

  • Smart Redaction
  • Microsoft Azure
  • GPT-5.6 (Terra)
Person names
0.93
0.87
0.91
Company names
0.81
0.70
0.67
Addresses
0.88
0.81
0.78
00.250.500.751.00
The real test of a redaction tool is the fields that can’t be caught by pattern alone, names, companies and addresses, where it has to read the context. Smart Redaction beats both Azure and GPT-5.6 on all three, widest of all on company names (0.81 against 0.70 and 0.67). This is where a model built for redaction pulls ahead of a general service and a frontier LLM alike.

On person names, Smart Redaction scores 0.93 ahead of GPT-5.6 (0.91) and Azure (0.87). On addresses it reaches 0.88 against 0.78 and 0.81. The widest gap is on company names where Smart Redaction scores 0.81 while GPT-5.6 and Azure sit at 0.67 and 0.70. These are the fields most likely to leak in a real document, and the ones a model trained for the task handles better than either Azure or GPT-5.6

Why we publish these results

Most data room providers describe their redaction as accurate and leave it there. We would rather show you the numbers, name the tools we benchmark against, and let you see how we compare. We bring the same openness to the rest of our Smart VDR suite too.

See how Smart Redaction performs on your own documents – Book a Demo

Footnotes

1. GPT-5.6 comes in more than one variant. Before testing, we asked GPT-5.6 Sol which to use for redaction, and it pointed us to Terra: redaction is a pattern-recognition task, which Terra is better at, and it needs none of the step-by-step reasoning Sol is built for. We ran Terra zero-shot, prompting it to find and label PII by category with no task-specific training, the way a team would out of the box.

2. The 80 documents spanned seven languages: English, Dutch, German, French, Polish, Italian and Czech. The blog highlights English and German as two widely used examples, and Polish and Czech as lower-resource ones; Smart Redaction led Azure in every language tested.

Want to find out more?

Submit your details below and a member of our team will get back to you shortly

This field is for validation purposes and should be left unchanged.