Webhook icon
Commands icon
Docker icon
OutputValues icon
Log icon
Switch icon
Set icon
If icon
SlackIncomingWebhook icon
Fail icon

Block a release when k6 shows it is slower than the last good one

Run k6 after each deploy with Kestra, compare p95 with the last good release in KV, block slow releases. Runs against Grafana's public demo app.

Categories
CoreInfrastructure

A fixed latency budget catches a release that is slow in absolute terms. It misses the release that went from 290 ms to 1,200 ms but is still "under 3 seconds". This flow load tests every release with k6, compares p95 latency with the last release that passed, and blocks the release when it got more than max_p95_regression_pct slower. A release that passes becomes the new baseline, so the bar follows the service.

It runs with no setup. The default target is QuickPizza, Grafana's public k6 demo application, with 3 virtual users for 20 seconds. Run it again with endpoint set to /api/delay/1 to watch a slow release get blocked.

This blueprint was created by zkasuran.

How it works

  1. load_test (io.kestra.plugin.scripts.shell.Commands, grafana/k6:2.3.0) runs a k6 script with two hard thresholds: p95 under max_p95_ms and an error rate under max_error_rate. Its handleSummary() prints p95, p99, median, error rate and request count as Kestra outputs and writes summary.json. k6 exits with code 99 when a threshold breaks. The task records the code and only fails on other codes.
  2. read_baseline (io.kestra.plugin.core.output.OutputValues) reads the last good release with the kv() function.
  3. verdict decides:
    • THRESHOLD_BROKEN: k6 exited with 99.
    • FIRST_RUN: no baseline yet.
    • REGRESSED: p95 is above baseline p95 * (1 + max_p95_regression_pct / 100).
    • PASS: otherwise.
  4. gate (io.kestra.plugin.core.flow.Switch): PASS and FIRST_RUN store this release as the baseline (io.kestra.plugin.core.kv.Set). The other verdicts keep the baseline, optionally post to Slack, and end the run as FAILED (io.kestra.plugin.core.execution.Fail), so the deploy pipeline can roll back.
  5. after_deploy (io.kestra.plugin.core.trigger.Webhook) lets the pipeline call the gate with {"release": "v1.4.2"} in the body.

Inputs

  • service (STRING, default quickpizza): names the baseline, k6_baseline_<service>.
  • base_url (STRING) and endpoint (STRING, default /api/quotes): what to load.
  • release (STRING, default v1.0.0): stored with the baseline. A webhook body release wins.
  • vus (INT, default 3) and duration (STRING, default 20s): keep them small against a shared demo target.
  • max_p95_ms (INT, default 3000) and max_error_rate (FLOAT, default 0.01): hard k6 thresholds.
  • max_p95_regression_pct (INT, default 25): allowed slowdown against the baseline.
  • notify_slack (BOOL, default false).

Prerequisites

  • A Kestra worker that can run Docker containers.

Secrets

  • SLACK_WEBHOOK_URL: only when notify_slack is true.

Quick start

  1. Run the flow. It reads FIRST_RUN and stores the baseline, about 290 ms p95.
  2. Run it with release set to v1.1.0. It reads PASS and moves the baseline.
  3. Run it with endpoint set to /api/delay/1. p95 jumps to about 1,280 ms, the run reads REGRESSED and fails, and the baseline is kept.
  4. Run it with endpoint set to /api/pizza (a 404). The error rate is 1, so k6 breaks its threshold and the run reads THRESHOLD_BROKEN.

Expected outputs

  • outputs.load_test.vars: p95_ms, p99_ms, median_ms, error_rate, requests, k6_exit.
  • outputs.verdict.values.state plus the baseline and the limit it was compared with.
  • outputs.load_test.outputFiles['summary.json']: the full k6 summary.
  • KV k6_baseline_<service>: {release, p95_ms, error_rate, execution_id}.

How to extend

  • Load several endpoints with k6 scenarios and keep one baseline per scenario.
  • Compare p99 or a custom k6 metric the same way.
  • Run k6 in the cloud or with more load generators by changing the command.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.