Schedule icon
Webhook icon
AIAgent icon
GoogleGemini icon
Script icon
Process icon
If icon
SlackIncomingWebhook icon

AI Incident Post-Mortem and Root Cause Generator

Automate blameless incident post-mortems with Gemini AI. Ingest alert timelines and error logs to extract 5 Whys, root causes, and action items.

Categories
AICoreInfrastructure

After high-severity production outages, engineering teams face significant friction drafting blameless post-mortem reports. On-call engineers are exhausted, timeline entries are scattered across monitoring tools (Datadog, Grafana, PagerDuty) and chaotic Slack war rooms, and action items often lack specific ownership. Consequently, post-mortems are delayed for weeks or skipped entirely, allowing repeat failure modes to cause future outages.

This blueprint automates blameless post-mortem generation. Once an incident is marked resolved in PagerDuty or Opsgenie, webhook integration triggers Kestra. An AI agent powered by Google Gemini analyzes raw chronological logs, establishes the sequence of events, calculates Time to Detect (TTD) and Time to Resolve (TTR), conducts a rigorous blameless "5 Whys" root-cause breakdown, and outlines concrete engineering remediation tasks.

The flow formats the analysis into an incident-postmortem.md markdown artifact saved to execution storage. For SEV-0 and SEV-1 incidents, an automated gate routes the executive summary card directly to engineering leadership in Slack to ensure high-priority remediation.

How it works

  1. analyze_incident_timeline (ai.agent.AIAgent): Ingests the chronological incident timeline and diagnostic observations, using gemini-2.5-flash with a strict JSON schema to construct the 5 Whys failure chain, compute TTD/TTR metrics, and assign preventative action items.
  2. render_postmortem_markdown (scripts.python.Script): Transforms the validated JSON output into a formatted incident-postmortem.md report saved directly in Kestra execution storage.
  3. severity_routing_gate (core.flow.If):
    • High Severity (SEV-0 / SEV-1): Delivers an executive briefing card to #engineering-leadership in Slack detailing root cause and mitigation tasks.
    • Standard Severity (SEV-2 / SEV-3): Delivers a notification to #devops-incidents for team-level tracking.
  4. alert_on_failure (errors block): Catches unexpected script or model failures and notifies the SRE operations team.
  5. Triggers: Weekly post-mortem sweep schedule (disabled: true by default) plus an authenticated Webhook trigger for incident resolution webhooks.

What you get

  • Sub-minute blameless post-mortem draft generation immediately upon incident resolution.
  • Rigorous 5 Whys root cause analysis eliminating superficial blame.
  • Standardized markdown post-mortem documentation archived in Kestra storage.

Who it's for

  • Site Reliability Engineers (SREs), DevOps, and Platform Engineering teams.
  • Engineering Managers and Directors seeking visibility into production stability.
  • Incident Commanders responsible for post-incident review facilitation.

Why orchestrate this with Kestra

Building incident summarizers as ad-hoc scripts or bot integrations often fails when logs are incomplete and lacks reliable error handling. Kestra provides declarative DAG orchestration, secure secret management, artifact persistence, and automatic routing gates for severity escalation.

Prerequisites

  • Google Gemini API key with access to gemini-2.5-flash.
  • Slack incoming webhook endpoint for incident reporting.

Secrets

  • GEMINI_API_KEY: API key for Google Gemini model inference.
  • SLACK_WEBHOOK_URL: Slack Incoming Webhook URL for incident notifications.
  • WEBHOOK_KEY: Authentication secret for event-driven webhook invocation.

Quick start

  1. Set GEMINI_API_KEY, SLACK_WEBHOOK_URL, and WEBHOOK_KEY in your Kestra namespace secrets.
  2. Click Execute in the Kestra UI to run the blueprint against the default sample database lock incident.
  3. Download the generated incident-postmortem.md from the Outputs tab and review the Slack card.

How to extend

  • Add tasks before analyze_incident_timeline to pull chat logs directly from Slack channels (via io.kestra.plugin.core.http.Request) or metrics snapshots from Datadog/Grafana.
  • Connect Jira or Linear plugins inside severity_routing_gate to automatically file preventive action items as tracked tickets in your engineering backlog.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.