Webhook icon
Schedule icon
Commands icon
Docker icon
If icon
SlackIncomingWebhook icon
Log icon

Verify a DNS Change Across Independent Resolvers Before You Celebrate

Poll a changed DNS record across independent public resolvers until every answer matches your target, and alert Slack listing stale resolvers on timeout.

Categories
CloudInfrastructure

DNS changes are the rare deploy step where "it worked on my machine" is literally true — your resolver cached the new answer minutes ago, the rest of the world is still on the old one, and the rollout either hangs or quietly fails for a subset of users. This blueprint turns propagation from a question you ask a website into a verification your pipeline executes: independent public resolvers are polled on an interval until every answer matches the value you set, and on timeout the alert lists exactly which resolvers were stale and what they returned.

How it works

  1. check_propagation (io.kestra.plugin.scripts.shell.Commands on the Docker task runner) installs dig in a throwaway alpine:3.20 container and loops: every resolver in the resolvers input is queried independently (dig @resolver name type), answers are compared against expected_value, and the round's match count is logged. When all resolvers agree, it emits all_matched: true with the round count; when the timeout_seconds deadline passes first, it emits all_matched: false plus a sanitized pending list of resolver=value pairs still disagreeing. An empty resolver list exits with an error instead of reporting a vacuous success.
  2. check_result (io.kestra.plugin.core.flow.If) branches on all_matched. False: alert_stale_resolvers posts the matched count, rounds, and the stale resolver/value pairs to Slack. True: log_propagated records how many rounds full consensus took.
  3. The errors block alerts Slack when the verifier itself fails — no dig, no network — because a crashed check must never be read as "propagated".
  4. Triggers: a Webhook (dns-propagation-check) for the deploy pipeline to call after terraform apply or the console change, plus an optional disabled Schedule for drift watching.

What you get

  • A yes/no propagation verdict your pipeline can branch on, not a browser refresh loop.
  • Independent vantage points: four different operators' resolvers, not four queries to your own cache.
  • On timeout, the exact resolver=value pairs that are still stale — proof propagation is incomplete rather than the check being wrong.
  • A propagation_summary JSON output for the next pipeline stage or a dashboard.

Who it's for

  • DevOps engineers cut over domains, migrations, and blue/green endpoints and need certainty before flipping the next switch.
  • Platform teams whose CD pipelines currently sleep 30 minutes and hope.
  • Anyone who has been burned by "propagated" meaning "propagated for me".

Why orchestrate this with Kestra

The polling loop itself is trivial shell; what Kestra adds is everything around it: an event entry point the deploy pipeline can call, a branch that separates consensus from timeout, an alert that carries the evidence, self-reporting failures, and an execution history that shows how long propagation actually took each time. Turned into a flow, the same check doubles as a drift detector by enabling the schedule — one blueprint, cutover verification and ongoing truth.

Prerequisites

  • Outbound UDP/53 from the Kestra Worker's Docker daemon to the resolvers you list (corporate networks sometimes block it — a good reason to list resolvers you can reach).
  • The expected value after your change: target IP for A/AAAA, target hostname for CNAME, full string for TXT.
  • A Slack incoming webhook for timeout and failure alerts.

Secrets

  • SLACK_WEBHOOK_URL: Slack incoming webhook URL for stale-resolver and verifier-failure alerts.

Quick start

  1. Add the Slack webhook secret to your namespace.
  2. Set hostname, record_type, and expected_value to the change you just made.
  3. Call the dns-propagation-check webhook from your pipeline — or run the flow manually — and read propagation_summary.
  4. Optionally enable drift_schedule to keep watching the record.

How to extend

  • Add the record's previous value as an input and alert when some resolvers return old and some new (split-brain) instead of only on mismatches.
  • Chain a second stage after check_result that runs smoke tests against the new target only once propagation is confirmed.
  • Verify many records from one change by wrapping the flow as a subflow over a records input.
  • Forward the timeout alert to PagerDuty by adding a notification task beside Slack.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.