Schedule icon
Webhook icon
Script icon
Process icon
Query icon
If icon
Log icon
SlackIncomingWebhook icon

AI Synthetic Relational Data Generator and Referential Integrity Gate

Generate realistic multi-table synthetic test datasets without PII. Verify foreign-key constraints and temporal consistency in DuckDB before QA staging.

Categories
AICoreData

Developing and testing mission-critical data pipelines requires rich, realistic datasets. However, dumping production customer data into staging or testing environments introduces catastrophic privacy risks, violating GDPR, HIPAA, and CCPA mandates.

Conversely, naive mocking scripts usually generate detached tables in silos: mock orders reference non-existent customer IDs, or payment timestamps precede order creation dates. When QA engineers run integration tests against inconsistent data, pipelines break or produce false test signals.

This blueprint generates realistic, multi-table synthetic test datasets (users, orders, payments) with zero personally identifiable information (PII). Ingested synthetic tables are loaded into an embedded in-memory DuckDB engine to verify foreign-key referential integrity and consistency across table relationships. Clean datasets are logged and published to Kestra execution storage for downstream staging databases, while any detected integrity violations immediately halt pipeline progression.

How it works

  1. generate_relational_entities (scripts.python.Script): Synthesizes realistic multi-table relational data using strictly randomized pseudonyms and identifiers, writing normalized CSV tables (synthetic_users.csv, synthetic_orders.csv, synthetic_payments.csv).
  2. validate_referential_integrity (jdbc.duckdb.Query): Executes an embedded DuckDB query with CTE joins to audit foreign-key relationships (orders.user_id -> users.user_id and payments.order_id -> orders.order_id), returning violation metrics and a boolean validity flag.
  3. data_quality_gate (core.flow.If):
    • Clean State (Zero Violations): Logs dataset metrics and alerts #data-quality-alerts in Slack that the dataset is verified and ready for staging ingestion.
    • Violation State (>0 Violations): Halts staging deployment and alerts engineers of foreign-key constraint failures.
  4. alert_on_failure (errors block): Catches unexpected script execution errors and notifies the data operations channel.
  5. Triggers: Scheduled weekly data refresh (disabled: true by default) plus an authenticated Webhook trigger for CI/CD staging pipelines.

What you get

  • 100% privacy-compliant test datasets with zero PII exposure risks.
  • Guaranteed multi-table referential integrity validated via in-memory DuckDB queries.
  • Production-grade quality gates preventing malformed mock data from polluting QA environments.

Who it's for

  • Data Engineers and Platform Engineers maintaining staging and pre-production databases.
  • QA Automation and SDET teams requiring realistic relational test fixtures.
  • Security and Compliance teams enforcing zero-production-data-in-staging mandates.

Why orchestrate this with Kestra

Maintaining custom seeding shell scripts across local environments leads to inconsistent test states and unverified data corruption. Kestra provides declarative DAG orchestration, embedded in-memory SQL execution with DuckDB, artifact lineage tracking, and automated Slack notification gates.

Prerequisites

  • Kestra instance with Python script runner capability.
  • Slack incoming webhook endpoint for data quality notifications.

Secrets

  • SLACK_WEBHOOK_URL: Slack Incoming Webhook endpoint for data quality alerts.
  • WEBHOOK_KEY: Authentication secret for event-driven webhook invocation.

Quick start

  1. Set SLACK_WEBHOOK_URL and WEBHOOK_KEY in your Kestra namespace secrets.
  2. Click Execute in the Kestra UI to run the generator with default input counts.
  3. Download the generated CSV artifacts from the execution Outputs tab.

How to extend

  • Connect downstream database tasks (such as io.kestra.plugin.jdbc.postgresql.Query or io.kestra.plugin.gcp.bigquery.Load) inside the approved branch to load validated synthetic tables directly into staging schemas.
  • Add an LLM task (io.kestra.plugin.ai.agent.AIAgent) to synthesize context-aware company domains, product descriptions, or support feedback categories.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.