Schedule icon
Script icon
Docker icon
Log icon
If icon
Queries icon
Query icon
UploadFiles icon
Fail icon

Measure a Presidio PII scrub, then block the load when residual PII survives

Scrub free text with Presidio in Kestra, count what each recognizer removed, re-scan with an independent detector, then block the load on any survivor.

Categories
AIData

Redaction steps in production are trusted on faith. A team runs a scrubber, sees no error, loads the batch, then finds a customer name in an analytics table months later. Nothing in that pipeline ever measured what was removed or checked whether anything survived.

This flow measures both. Presidio scrubs the free-text column, the per-entity redaction counts are published as an artifact, then the scrubbed output is re-scanned by a second detector that shares no code with the first. The warehouse load sits inside an If branch, so residual PII has no path into the table rather than a warning beside it.

This blueprint was created by zkasuran.

Why the second detector is not Presidio

Re-running the same AnalyzerEngine over anonymised text returns almost nothing, because the text no longer contains what that engine was built to match. It is a check that cannot fail. The flow measures that number anyway (same_detector_rescan_hits) so the report shows what it is worth, then makes the verdict on a separate battery of closed-form validators:

  • Card numbers: every digit run of 13 to 19 digits, separators stripped, tested with the Luhn double-add-double checksum.
  • IBANs: the first four characters moved to the end, letters mapped A=10 through Z=35, then checked for a remainder of 1 modulo 97.
  • Phone numbers: scanned with phonenumbers, which presidio-analyzer already depends on, at Leniency.VALID, capped at the 15 digits ITU-T E.164 allows.
  • Mailboxes: one strict address pattern.
  • Your own identifiers: a list of regular expressions in the org_id_patterns input, for employee numbers, case references, internal account ids. Presidio ships no recognizer for these, so nothing else in the pipeline is looking.

The report names which residual types Presidio never flagged in the first pass (types_presidio_never_flagged). On the default run that list is not empty, which is the whole argument for the second pass.

Residual values are written out masked, never in full, so the findings artifact is safe to hand to a reviewer who is not cleared for the source data.

How it works

  1. generate_batch (io.kestra.plugin.scripts.python.Script) writes a synthetic ticket batch to tickets.csv. Every value is invented. With inject_unsupported_ids on it also writes an employee id, an internal account reference plus a dot-separated card number.
  2. scrub_and_rescan installs presidio-analyzer and presidio-anonymizer pinned to 2.2.364, downloads the spaCy model named by the spacy_model input, then runs AnalyzerEngine.analyze followed by AnonymizerEngine.anonymize per ticket. It counts every replacement by entity type, re-runs the analyzer over its own output for comparison, runs the independent battery, then emits scrubbed.csv, redaction-report.json plus residual-findings.json.
  3. log_measurement prints the counts, the hollow-check number and the residual verdict in one line.
  4. residual_gate (io.kestra.plugin.core.flow.If) compares the residual count against max_residual_entities.
    • Clear: load_warehouse (io.kestra.plugin.jdbc.duckdb.Queries) inserts the scrubbed rows stamped with the execution id, verify_load (io.kestra.plugin.jdbc.duckdb.Query) reads them back, publish_report stores the batch with its measurement under pii-gate/passed/.
    • Not clear: quarantine_batch stores the batch and the masked findings under pii-gate/blocked/, then block_load (io.kestra.plugin.core.execution.Fail) ends the run as FAILED naming the count, the threshold plus the surviving types. The warehouse is untouched.

Honest limits

  • en_core_web_sm is the default because it is 12.8 MB against 400.7 MB for en_core_web_lg, which keeps a cold start near a minute and a half. Its PERSON recall is materially worse. Switch the spacy_model input to en_core_web_lg in production, since a name the model misses is a name the residual battery cannot catch either.
  • The residual battery finds what has a checksum or a pattern. Names, addresses, free-text medical detail or anything else without a closed form stay the model's job. The gate raises the floor, it does not replace review.
  • warehouse_path defaults to /tmp, which is wiped with the container. Point it at a mounted volume to keep history.

Prerequisites

  • A Kestra worker that can run Docker task runners, for the two Python script tasks.
  • Outbound access to PyPI plus the spaCy model release on first run, about 105 MB with the small model.
  • A writable path for the DuckDB file if you want the loaded rows to persist.

No secrets are needed. The flow is self-contained and safe to run as-is.

Quick start

  1. Execute the flow with the defaults. It blocks, because the batch carries identifiers Presidio has no recognizer for.
  2. Read the block_load error and the log_measurement line: the per-entity redaction counts, the residual count, the surviving types plus the types Presidio never flagged.
  3. Open pii-gate/blocked/<execution>/residual-findings.json in the namespace files. Values are masked.
  4. Re-run with inject_unsupported_ids set to false. The batch scrubs clean, the residual count is 0, DuckDB takes the rows, verify_load reads them back.
  5. Point the first task at your own extract, then set org_id_patterns to the identifier shapes your organisation actually uses.

Expected outputs

  • scrubbed.csv: the batch with the free-text column anonymised plus a per-row count of removed spans.
  • redaction-report.json: the model used, the per-entity redaction counts, Presidio's first-pass entities, same_detector_rescan_hits, the independent residual count and types, plus types_presidio_never_flagged.
  • residual-findings.json: one masked record per survivor with its ticket id, detector and entity type.
  • outputs.scrub_and_rescan.vars: residual_entities, residual_types, missed_by_presidio, redacted_total, redaction_summary, same_detector_rescan_hits.
  • On the clear branch, outputs.verify_load.row.loaded_rows read back from DuckDB.

How to extend

  • Add validators to the battery for the identifiers you have checksums for: NHS numbers, social insurance numbers, VAT numbers, ISINs. One function plus one entry in the findings list.
  • Send redaction-report.json to your metrics backend and alert on a drop in redaction counts. A scrubber that quietly stops matching looks exactly like a batch with no PII in it.
  • Swap generate_batch for a database query, an S3 download or a webhook payload, then keep the rest unchanged.
  • Set max_residual_entities above 0 only with a documented reason, since it is the one input that lets known survivors through.
  • Replace AnonymizerEngine default replacement with OperatorConfig("hash") or OperatorConfig("encrypt") when downstream needs to join on the redacted value.
  • Run the analyzer from the mcr.microsoft.com/presidio-analyzer image, which bakes en_core_web_lg in, when you want the larger model without the download. It runs as user 1001 and ships no anonymizer, so it needs its own task rather than a one-line swap.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.