Webhook icon
Schedule icon
Query icon
If icon
Patch icon
Log icon
Parallel icon
LoopUntil icon
Script icon
SlackIncomingWebhook icon

Kubernetes Chaos Game Day with Parallel Blast-Radius Watch

Controlled pod-loss game day - SLO baseline, scale-down fault, parallel error-rate and readiness watch, restore, and a Slack evidence report.

Categories
Infrastructure

How it works

  1. game_day (event-based Webhook) starts the drill on demand, or weekly_drill (Schedule, shipped disabled) runs it every Saturday morning.
  2. capture_baseline (prometheus.Query, FETCH_ONE) records the pre-fault 5xx ratio for the target namespace - the number the evidence report compares against.
  3. inject_gate (core.flow.If) skips injection on dry_run rehearsals; otherwise scale_down (kubernetes.kubectl.Patch, JSON_MERGE) drops the deployment from replicas to replicas - 1, the abrupt pod loss the system must absorb.
  4. blast_radius_watch (core.flow.Parallel) runs both sides of the blast radius concurrently: watch_error_rate (core.flow.LoopUntil) polls the 5xx ratio every 15 seconds until it is back inside error_budget_pct, while watch_ready_replicas (LoopUntil) polls kube_deployment_status_replicas_ready until the deployment is serving at its degraded size. Both loops share max_watch as a ceiling with failOnMaxReached: true, so a system that never stabilizes fails the drill instead of hanging.
  5. restore_capacity (kubectl.Patch) returns the deployment to its steady-state replica count, then sample_recovery and sample_ready capture post-restore numbers as top-level outputs.
  6. write_evidence (scripts.python.Script, Process runner) computes the verdict - STABLE (recovery within budget and replicas restored), DEGRADED (replicas back but errors over budget), or NOT_RESTORED - plus headroom and the fault-injected flag, via Kestra's ::json:: outputs protocol.
  7. announce_evidence posts the summary to Slack with messageText; the errors block posts a distinct alert when the watch hits maxDuration or the patch is rejected.

What you get

  • A repeatable, scheduled drill instead of an ad-hoc kubectl scale during someone's lunch break.
  • Parallel client-side and server-side watches, so "it stabilized" means both error rate and ready replicas - not just one of them.
  • A hard time ceiling that turns a never-recovering system into a red execution with an alert.
  • An evidence record: baseline, recovery, headroom, verdict, and execution id, in Kestra and in Slack.

Who it's for

SRE and platform teams practicing game days on Kubernetes workloads who want drills that leave a paper trail, and teams preparing for incident exercises where reviewers ask "how did you know it recovered".

Why orchestrate this with Kestra

The usual drill is a shared shell script plus eyeballing Grafana. Kestra gives the drill typed inputs, parallel watches with their own timeouts, durable state while someone observes, and a verdict that becomes a queryable output - and when the drill fails, that failure is an execution with logs, not a scrollback in a terminal.

Prerequisites

  • A Kubernetes deployment (default inputs expect 3 replicas in namespace chaos-lab) serving traffic that Prometheus scrapes with http_requests_total labeled by namespace.
  • Prometheus exporting kube_deployment_status_replicas_ready (kube-state-metrics).
  • A rehearsal with dry_run: true before injecting real faults.

Secrets

  • PROMETHEUS_URL: Prometheus base URL, e.g. http://prometheus:9090.
  • K8S_MASTER_URL: Kubernetes API server URL.
  • K8S_TOKEN: service account token with deployment patch permissions.
  • SLACK_WEBHOOK_URL: Slack incoming webhook for evidence and failure alerts.

Quick start

  1. Set the four secrets in your Kestra namespace.
  2. Adjust target_namespace, target_deployment and replicas to your workload.
  3. Run once with dry_run: true to confirm baseline, watches, and Slack wiring - no fault is injected.
  4. Fire the real drill via the game_day webhook and watch the Parallel branches in the execution graph.
  5. Read verdict, headroom and fault_injected in the outputs; enable weekly_drill once you trust the setup.

How to extend

  • Add a third Parallel branch that watches a business metric (orders/min) alongside errors and readiness.
  • Replace the scale-down fault with kubectl.Delete of a single pod once you pass a victim pod name from a kubectl.Get.
  • Feed verdict into a follow-up flow that opens an incident or a tracking issue when the drill fails.
  • Vary the fault: patch CPU limits, drain a node, or block egress with a NetworkPolicy - only the inject task changes.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.