Schedule icon
Webhook icon
Commands icon
Docker icon
If icon
Script icon
SlackIncomingWebhook icon
Log icon
OutputValues icon

Detect Kubernetes CrashLoops and Post Triage Digests to Slack

Kubernetes CrashLoop monitor with Kestra. Detect backoff pods, capture restart history, and alert Slack with triage steps.

Categories
Infrastructure

Find crashing pods before they page you at 3am. This blueprint scans your cluster every 15 minutes for CrashLoopBackOff, ImagePullBackOff, ErrImagePull, and pods above a restart threshold, builds a triage digest with per-pod kubectl logs and describe remediation hints, and posts it to Slack. Healthy runs log an all-clear instead of spamming the channel.

How it works

  1. every_15_minutes (Schedule, */15 * * * *) launches the scan; on_demand_webhook allows ChatOps triage.
  2. list_unhealthy_pods (shell Commands on bitnami/kubectl) dumps kubectl get pods -A -o json and filters in Python to unhealthy.json.
  3. build_triage_report (Python Script) formats the digest with restart counts and copy-paste kubectl logs --tail commands.
  4. decide_alert (If on breached) posts via Slack webhook or logs healthy.
  5. emit_triage_metrics (OutputValues) exports counts for dashboards.

What you get

  • Cluster-wide CrashLoop and image-pull detection in one flow.
  • Restart-threshold tuning (default 5) plus namespace filtering.
  • Actionable Slack digest with exact kubectl commands.
  • Quiet on healthy runs; loud on real incidents.

Who it's for

  • Platform and SRE teams running multi-namespace Kubernetes.
  • Teams without Prometheus/Kube-state-metrics wanting lightweight detection.
  • On-call rotations needing fast triage context.

Prerequisites

  • Kestra worker with Docker runner and cluster network access.
  • kubeconfig or in-cluster RBAC permitting get pods, logs, describe.
  • Slack Incoming Webhook URL.

Secrets

  • SLACK_WEBHOOK_URL: Slack destination for triage digests.
  • WEBHOOK_KEY: auth key for the on-demand webhook trigger.
  • Kube credentials per your runner setup (kubeconfig mount or service account).

Quick start

  1. Set secrets and grant pod-read RBAC.
  2. Set namespaces to all or a comma list; tune restart_threshold.
  3. Run manually to verify pod listing works.
  4. Enable the every_15_minutes schedule.

How to extend

  • Add a Parallel branch that captures kubectl logs --tail per pod into outputs.
  • Auto-restart Deployments stuck in ImagePullBackOff after registry credential rotation.
  • Write every triage to Postgres for MTTR reporting.

Links

Inputs

  • namespaces (STRING, default all): comma list or all.
  • restart_threshold (INT, default 5): flag at or above this count.
  • log_tail_lines (INT, default 100): log lines referenced per pod.

Outputs

  • outputs.build_triage_report.vars.count: unhealthy pod count.
  • emit_triage_metrics.values.triaged_at: scan timestamp.
See How

New to Kestra?

Use blueprints to kickstart your first workflows.