Schedule icon
Webhook icon
Script icon
If icon
SlackIncomingWebhook icon
Log icon

Probe Endpoint Latency and Gate on p95

Synthetic latency probe matrix that alerts Slack when p95 exceeds your budget or requests fail.

Categories
CoreInfrastructure

Uptime says "up"; p95 says "usable". This blueprint fires N timed requests at your URL every 15 minutes, computes min/mean/p95 plus failure count, and pages Slack only when the p95 budget breaks. Clean runs log their numbers so your execution history is the latency trend.

How it works

  1. probe_latency (io.kestra.plugin.scripts.python.Script) sends samples requests, times each, treats errors as 10s sentinels, and emits p95, mean, min, failures via the ::{"outputs": ...}:: protocol.
  2. check_budget (io.kestra.plugin.core.flow.If) branches: breach → alert_latency_breach with the stats; healthy → log_within_budget.
  3. The errors block alerts Slack when the probe script itself fails.
  4. Triggers: a 15-minute Schedule (disabled by default) plus a Webhook for incident-time probes.

What you get

  • A p95-based alert instead of "is it down".
  • Failure counts alongside latency, not instead of them.
  • latency_summary JSON for dashboards or SLO calcs.

Who it's for

  • Teams without a paid synthetic monitoring tool.
  • On-call engineers who want a quick incident probe.
  • SRE teams prototyping an SLO before buying one.

Why orchestrate this with Kestra

A curl loop in cron gives you numbers nowhere to go. The flow adds the schedule, the budget branch, the Slack alert, the failure channel, and history — and the next step (region matrix, status page, incident ticket) is one task away.

Prerequisites

  • The URL reachable from the Kestra Worker.
  • A Slack webhook.

Secrets

  • SLACK_WEBHOOK_URL: webhook for breach and probe-failure alerts.

Quick start

  1. Add the Slack webhook secret.
  2. Set url, samples, and p95_budget_ms.
  3. Run once and read latency_summary.
  4. Enable every_15min.

How to extend

  • Loop over several URLs with a ForEach.
  • Store latency_summary in KV and alert only on spikes vs. baseline.
  • Add a check_cert-style companion probe on the same schedule.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.