Parallel icon
Request icon
If icon
SlackIncomingWebhook icon
Log icon
Schedule icon

Monitor API Uptime in Parallel and Alert Slack with the Exact Status of Each Endpoint

Check several endpoints in parallel on a schedule with Kestra, treat any non-200 as unhealthy, and alert Slack with the status code of every site.

Categories
CoreInfrastructureinfrastructure

The simplest monitor most teams never get around to building: hit a handful of endpoints on a schedule and shout if any of them answer with something other than HTTP 200. This blueprint probes three sites in parallel every thirty minutes and sends one Slack message that names every endpoint together with the status it actually returned, so the person reading the alert knows which site to look at before they open anything.

The one detail that makes it work is options.allowFailed: true on each probe. Without it the HTTP task fails outright on any non-2xx response, the flow jumps straight to the errors block, and you lose the status code you wanted to report. With it, a 503 is a result to evaluate rather than a crash.

How it works

  1. The every_30_minutes trigger (io.kestra.plugin.core.trigger.Schedule) starts a run on the half hour.
  2. The check_all_sites task (io.kestra.plugin.core.flow.Parallel) fires three io.kestra.plugin.core.http.Request probes at once, each with a thirty second timeout and options.allowFailed: true, so one slow endpoint neither delays nor masks the others.
  3. The evaluate_results task (io.kestra.plugin.core.flow.If) reads outputs.<probe>.code for each site and treats anything other than 200 as unhealthy.
  4. The then branch posts a Slack message listing every site and its status code. The else branch logs a one line all clear so the execution history reads cleanly.
  5. The errors block sends a different message for the case where a probe could not run at all, such as a DNS failure or a hard timeout, which is a different kind of problem from a site returning a bad status.

What you get

  • One alert per failed run that names each endpoint and the status it returned, instead of a generic "something is down".
  • Parallel probing, so the schedule interval is bounded by your slowest site rather than the sum of all of them.
  • A clear split between "a site answered badly" and "we could not reach a site", routed to two different messages.
  • A run history with the status code of every probe on every run, which is a free uptime log.

Who it's for

  • Teams that want basic uptime visibility without buying a monitoring product for three URLs.
  • Developers who own a few public endpoints and want to know before their users do.
  • Anyone who already runs Kestra and would rather add a flow than another service.

Why orchestrate this with Kestra

A cron script that curls three URLs and pipes to a webhook is thirty lines, and it loses the thing that matters: history. Kestra keeps every probe's status code on every run, shows the three checks side by side in the Gantt view, retries the Slack send if the webhook hiccups, and gives you a second alert path for the case where the probe itself breaks. When a site has been flapping for a week, the execution list already contains the evidence.

Prerequisites

  • A Slack incoming webhook for the alert channel.
  • Outbound HTTPS access from the Kestra worker to the sites you want to probe.

Secrets

  • SLACK_WEBHOOK: incoming webhook URL used for both the unhealthy alert and the probe failure alert.

Quick start

  1. Add the SLACK_WEBHOOK secret to your Kestra namespace.
  2. Replace the three site_* input defaults with your own endpoints, or leave them to see the flow run against public sites first.
  3. Execute once manually and check the evaluate_results output to confirm you see a status code per site.
  4. Enable the trigger. Adjust the cron if thirty minutes is too coarse or too chatty for you.

How to extend

  • Monitor more endpoints by adding further check_site_* tasks under the Parallel block, or switch to io.kestra.plugin.core.flow.Loop over an ARRAY input when the list grows.
  • Add a latency budget by comparing each probe's duration in the condition, so a site that answers 200 in nine seconds still counts as unhealthy.
  • Replace the Slack message with io.kestra.plugin.pagerduty.alerts.Create for endpoints that should page someone.
  • Require two consecutive failures before alerting by reading the previous execution's state with the Kestra KV store, which filters out single blips.
  • Post the status codes to a dashboard or a metrics endpoint with another http.Request to chart uptime over time.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.