OutputValues icon
Queries icon
Log icon
UploadFiles icon
If icon
Fail icon
Set icon
Schedule icon

Report the four DORA metrics per service every week and alert when delivery degrades

DORA metrics in Kestra and DuckDB. Deploy frequency, lead time, change failure rate, time to restore, performance bands and degradation alerts.

Categories
DataInfrastructure

The four DORA metrics are the standard way to measure software delivery: how often a team ships, how long a change takes to reach production, how often a change breaks production, and how fast service is restored. Most teams compute them in a vendor tool or once a year by hand. The data already exists in the CI/CD log, Git and the incident tool. The hard parts are joining them per deployment, and noticing when a team starts sliding.

This flow joins them every Monday, computes the four metrics per service over a rolling window, places each service in a DORA performance band and compares it with the previous report over the same window.

It runs with no setup and no network. The demo has six weeks of history for four services:

  • checkout ships several times a day;
  • ledger ships every two weeks;
  • search started batching its changes, so its lead time grows;
  • mobile-api has a bad fortnight, with three failed changes and one long outage.

How the metrics are computed

Metric Rule
Deployment frequency Successful production deploys per week over the window. A deploy that failed in the pipeline never reached users and is not counted
Lead time for changes Median time from commit to the production deploy that shipped it, over every commit in the window
Change failure rate Deploys that caused an incident (rollback, hotfix, linked incident), divided by deploys
Time to restore Median time from incident start to resolution

Bands

The thresholds follow the Accelerate State of DevOps report. The overall band is the weakest of the four metrics, so a team that ships hourly but breaks production every other deploy is not called elite.

Band Deploys Lead time Change failure rate Restore
ELITE 7 or more a week Under a day 5% or less Under an hour
HIGH 1 a week Under a week 10% or less Under a day
MEDIUM 1 a month Under a month 15% or less Under a week
LOW Less Longer Higher Longer

A metric is flagged as degraded when it is worse than the previous report over the same window by more than degrade_pct. The change failure rate must also move by at least 2 points, so one incident on a quiet service is not an alert.

How it works

  1. state loads the latest earlier report over the same window from KV.
  2. metrics (io.kestra.plugin.jdbc.duckdb.Queries) holds the services, deployments, commits and incidents. It runs the checks, then computes the metrics, the bands and the comparison.
  3. log_report prints every service. evidence_blockers stores blockers.csv.
  4. gate stops on blockers, before any report or history is written. Otherwise evidence stores services.csv (metrics, band, previous values, what degraded) and failed-changes.csv.
  5. save (io.kestra.plugin.core.kv.Set) adds this report to the history, keyed by window and week.
  6. degraded_gate logs a warning with what degraded.

Blockers: a successful deploy with no commit linked (it has no lead time and hides failures), an incident resolved before it started, and a deploy for a service not in the catalog.

Tested end to end

On Kestra 2.0.5 OSS, window_days = 14.

Run Week ending Result
1 2026-09-20 checkout HIGH (24 a week, 1.4h, 6.3%), ledger MEDIUM (0.5 a week), mobile-api and search HIGH. No history, nothing flagged
2 2026-09-27 mobile-api LOW: change failure rate 0% to 50%, restore 11.29h, lead time 41.2h to 52.0h. search lead time 23.5h to 31.9h, change failure rate 0% to 9.1%. Warning logged
3 2026-10-04 mobile-api change failure rate 50% to 75%. search lead time 31.9h to 62.3h. checkout 6.3% after 8.3%, not flagged
4 2026-10-04, 28 days Compared with nothing: the 14-day history is kept apart
5 ISSUES Blocked: two hotfix deploys with no commit, an incident resolved before it started

I recomputed the bands, the change failure rates, the deploy frequencies and the 6 degradation flags of the 12 service-weeks independently in Python: every value matches.

Inputs

Input Default Purpose
week_ending 2026-10-04 Last day of the report
window_days 28 Rolling window
degrade_pct 25.0 Degradation threshold
scenario CLEAN Demo only

Use it with your data

  • Replace deployments with your CD tool's production deploys (GitHub deployments, Argo CD, Kestra executions of your release flow), commits with the commits between two deploys, and incidents with your incident tool, linked to the deploy that caused them.
  • Enable the weekly trigger. It runs on Monday for the week ending the day before.
  • Send the warning from alert to the team channel, and services.csv to your dashboard.

Things to know

  • Lead time starts at the commit. Use the first commit of the pull request if you prefer to include review time.
  • A service with no incident in the window gets the top band for time to restore.

Links

This blueprint was created by zkasuran.

See How

New to Kestra?

Use blueprints to kickstart your first workflows.