Query icon
If icon
Apply icon
Request icon
Queries icon
SlackIncomingWebhook icon
Delete icon
Create icon
Fail icon
Schedule icon
Webhook icon

MLOps Canary Promotion Gate with Drift Watchdog

Promote ML models safely in Kestra: eval and drift thresholds gate a canary deploy, a smoke check verifies it, and failures roll back with an issue.

Categories
AI

Most model regressions reach production through a gap between "the evaluation looked fine" and "the deployment actually works." A notebook says the metrics passed, someone copies a version number into a manifest, and nothing checks that the new model can even answer a health request before traffic arrives. This blueprint closes that gap as one workflow: the candidate's evaluation and drift score decide whether a canary deploys at all, the canary has to prove it serves traffic before the registry flips to production, and every rejection or rollback leaves a trail in Slack and GitHub.

How it works

  1. load_candidate (io.kestra.plugin.jdbc.postgresql.Query, FETCH_ONE) pulls the newest row from model_evals with its latest drift score from model_drift via a LEFT JOIN LATERAL - one round trip gives the gate everything it needs.
  2. eval_gate (io.kestra.plugin.core.flow.If) compares accuracy, p95 latency, and drift against the threshold inputs. A rejection calls core.execution.Fail with the actual numbers in the message, so the flow-level errors handler delivers the reason to Slack and any upstream release pipeline stops.
  3. deploy_canary (io.kestra.plugin.kubernetes.kubectl.Apply) rolls the candidate out as a one-replica canary labeled with its model version; smoke_canary (io.kestra.plugin.core.http.Request with allowFailed) probes the health endpoint without aborting on a non-200.
  4. smoke_gate (io.kestra.plugin.core.flow.If) promotes only on a 200: demote_previous archives the old production version, promote_candidate upserts the new one (ON CONFLICT keeps the registry idempotent), and announce_promotion posts the numbers to Slack.
  5. A failed smoke deletes the canary (io.kestra.plugin.kubernetes.kubectl.Delete), opens a labeled rollback issue (io.kestra.plugin.github.issues.Create), and fails the execution - the registry is never touched.

What you get

  • A single gate for accuracy, latency, and drift thresholds - three JSON-free inputs, no SQL edits.
  • Canary-first promotion: nothing enters the registry without answering a live health probe.
  • Automatic rollback with an assignable GitHub issue carrying every metric.
  • promoted_model, accuracy, and drift_score outputs for dashboards and downstream release flows.
  • A nightly schedule or an event-driven webhook, so promotion fits either ops culture.

Who it's for

  • ML engineers who need promotion to be a pipeline, not a runbook performed at 11 p.m.
  • Platform teams operating model registries who want drift to veto stale candidates.
  • Organizations where release governance requires evidence (metrics, rollback, issue trail) for every production change.

Why orchestrate this with Kestra

Promotion scripts hide state: the notebook ran, the deploy happened, nobody knows which version gate failed. Kestra gives you explicit branches for each gate, execution history that records the metrics as outputs, a failure handler that notifies instead of exits silently, and native Kubernetes and GitHub plugins - so the gate, the deploy, and the evidence live in one reviewable workflow.

Prerequisites

  • PostgreSQL with three tables:

    CREATE TABLE model_evals (
      id BIGSERIAL PRIMARY KEY,
      model_id TEXT NOT NULL,
      version TEXT NOT NULL,
      status TEXT NOT NULL DEFAULT 'candidate',
      accuracy NUMERIC(6,4) NOT NULL,
      f1 NUMERIC(6,4) NOT NULL DEFAULT 0,
      p95_latency_ms INT NOT NULL DEFAULT 0,
      trained_at TIMESTAMPTZ NOT NULL DEFAULT now()
    );
    
    CREATE TABLE model_drift (
      id BIGSERIAL PRIMARY KEY,
      model_id TEXT NOT NULL,
      version TEXT NOT NULL,
      drift_score NUMERIC(6,4) NOT NULL,
      checked_at TIMESTAMPTZ NOT NULL DEFAULT now()
    );
    
    CREATE TABLE model_registry (
      model_id TEXT NOT NULL,
      version TEXT NOT NULL,
      status TEXT NOT NULL,
      promoted_at TIMESTAMPTZ,
      PRIMARY KEY (model_id, version)
    );
    
  • A Kubernetes cluster with your canary image, and a health endpoint reachable from Kestra.

  • A Kestra instance.

Secrets

  • POSTGRES_JDBC_URL: JDBC URL for the registry database.
  • POSTGRES_USERNAME: database username.
  • POSTGRES_PASSWORD: database password.
  • K8S_MASTER_URL: Kubernetes API server URL.
  • K8S_TOKEN: service-account token with apply/delete on the canary namespace.
  • SLACK_WEBHOOK_URL: Slack incoming webhook for promotion and failure messages.
  • GITHUB_TOKEN: token with issue-write access to the repository input.

Quick start

  1. Add the seven secrets above and create the three tables.
  2. Insert a candidate evaluation row and, optionally, a drift row for the same model and version.
  3. Set min_accuracy, max_p95_latency_ms, max_drift_score, canary_image, canary_url, and repository.
  4. Run via on_demand - watch load_candidate, the gate, the canary apply, and the smoke result.
  5. Break the smoke on purpose (point canary_url at a dead port) and confirm the delete, the GitHub issue, and the failed execution.
  6. Enable promotion_window when you trust the thresholds.

How to extend

  • Wait for readiness: add a LoopUntil around smoke_canary that retries for a minute before gating - cold models often need warm-up time.
  • Wire drift monitoring: a separate schedule that writes model_drift rows (from Prometheus or evidently) can chain into this flow with a Flow trigger, so fresh drift data immediately re-runs the gate.
  • Batch promotion: wrap the whole flow in a Subflow loop to promote several candidate models in one release train.
  • Progressive traffic: after smoke_gate, shift service traffic in steps (10% - 50% - 100%) with a Loop over kubectl apply weight patches before promote_candidate.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.