Script icon
Log icon
If icon
UploadFiles icon
Set icon
Fail icon
Schedule icon

Resolve customer duplicates across systems into golden records, with a quality gate

Deduplicate customers across systems with Splink in Kestra. Measure precision and recall, block over-merged entities, publish golden records.

Categories
AIData

The same customer lives in the CRM, billing and support with a typo here, a missing email there and a transposed letter somewhere else. Entity resolution links those records into one entity. Getting it wrong in either direction is costly: missed matches keep duplicates, and over-merging silently fuses different people (often relatives sharing a surname and an address) into one customer. This flow resolves records with Splink and publishes golden records only after it has checked both directions.

It needs no secrets or services, only network access to install Splink. A demo extract writes about 2,750 records for 1,200 people across three systems, including households where relatives share a surname, city and postcode. A quarter of the people form the clerically reviewed sample, records a person has already linked by hand. Run it with match_threshold: 0.9 to watch relatives get chained into one entity and the gate block it.

This blueprint was created by zkasuran.

How it works

  1. extract_sources (io.kestra.plugin.scripts.python.Script) writes crm.csv, billing.csv, support.csv and reviewed_sample.csv. Replace it with your real extracts.
  2. resolve_entities (io.kestra.plugin.scripts.python.Script, splink==5.0.0 on DuckDB):
    • blocks on surname, first name + birth date, email and postcode,
    • compares names with Jaro-Winkler and the other fields exactly,
    • trains the model without labels by expectation maximisation,
    • clusters records into entities at match_threshold,
    • measures pairwise precision and recall on the reviewed sample and finds the largest entity,
    • builds one golden record per entity (newest non-empty value, CRM wins a tie) and a crosswalk from every source record to its entity.
  3. log_quality reports the counts, the scores and the names in the largest entity.
  4. quality_gate (io.kestra.plugin.core.flow.If) checks precision against min_precision, recall against min_recall and the largest entity against max_cluster_size.
    • Pass: publish_golden_records (io.kestra.plugin.core.namespace.UploadFiles) writes mdm/customers/golden_records.csv and crosswalk.csv, and record_quality (io.kestra.plugin.core.kv.Set) stores the scores.
    • Fail: block_publication (io.kestra.plugin.core.execution.Fail) names the merged people. The previous golden records stay.
  5. nightly (io.kestra.plugin.core.trigger.Schedule).

Why both checks are needed

Threshold Entities Precision (reviewed) Recall Largest entity Verdict
0.99 (default) 1,216 1.0 0.984 6 Published
0.9 1,143 1.0 0.984 11: Barbara, Michael, Olu and Sarah Williams Blocked
0.5 983 1.0 0.996 16, the Kim household Blocked
0.1 942 0.952 0.996 16 Blocked

At 0.9 the reviewed sample still scores perfect precision, because the chained relatives fall outside it. The entity-size guard is what catches the over-merge. A sampled metric alone would have published four different people as one customer.

Inputs

  • match_threshold (FLOAT, default 0.99).
  • min_precision (FLOAT, default 0.98) and min_recall (FLOAT, default 0.95).
  • max_cluster_size (INT, default 6): records allowed in one entity.

Expected outputs

  • outputs.resolve_entities.vars: records, entities, multi_source_entities, largest_entity, largest_entity_names, reviewed_pairs, precision, recall.
  • Namespace files mdm/customers/golden_records.csv and mdm/customers/crosswalk.csv.
  • KV customer_entity_resolution_quality.

Using it with your own data

  • Replace extract_sources with your extracts, keeping a unique_id, a source and the matching columns.
  • Keep a reviewed sample of linked records from your data stewards. Even a few hundred pairs make precision and recall measurable.
  • Tune the blocking rules and comparisons to your columns. Splink's documentation covers the trade-offs.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.