Webhook icon
Schedule icon
Log icon
Script icon
SlackIncomingWebhook icon
If icon
Pause icon

Autonomous Incident Triage with Multi-Provider LLM Fallback and Slack Escalation

Triage production alerts using multi-tier LLM failover from OpenAI to Gemini, evaluate blast radius, pause for human approval, and dispatch Slack alerts.

Categories
AIBusinessInfrastructure

Automate production incident triage across mission-critical services with resilient multi-tier LLM failover. When production outages occur, on-call engineers face alert storms from PagerDuty, Sentry, Datadog, and cloud monitors. Single-model AI pipelines introduce a single point of failure (SPOF): if OpenAI experiences latency spikes, rate limits (HTTP 429), or outages, automated incident response stops dead in its tracks.

This workflow provides an enterprise-grade incident triage pipeline orchestrated by Kestra. It ingests alert events via authenticated webhooks or scheduled polling, triages diagnostic logs using OpenAI as the primary provider, automatically fails over to Google Gemini if OpenAI fails or times out, and falls back to a deterministic offline heuristic rule engine if external AI APIs are unreachable. High-severity incidents are held in a human-in-the-loop approval gate before running automated mitigation actions, keeping production safe and auditable.

Architecture and Failover Mechanics

The workflow executes a 3-tier failover hierarchy:

  1. Tier 1 (Primary - OpenAI): Attempts root cause analysis, blast radius estimation, and runbook generation using gpt-4o-mini with strict JSON schema constraints.
  2. Tier 2 (Secondary - Google Gemini): If OpenAI returns an HTTP error, times out after 12 seconds, or exhausts token quotas, execution instantly fails over to Google Gemini (gemini-1.5-flash) via the Gemini REST API.
  3. Tier 3 (Offline Deterministic Heuristic): If both external AI providers are unavailable (e.g. network partition or missing credentials), the flow executes an internal rule-based heuristic that maps log patterns (such as connection pool saturation or gateway timeouts) to remediation playbooks with requires_human_approval forced to true.
  4. Human-in-the-Loop Gate: High-severity or high-blast-radius incidents pause execution (io.kestra.plugin.core.flow.Pause) for up to 1 hour, notifying on-call responders via Slack. The execution only advances to mitigation runbooks when an engineer explicitly resumes the flow with an approval flag and change parameters.

How It Works

  1. Inbound Trigger: An alert payload arrives via the alert_webhook (io.kestra.plugin.core.trigger.Webhook) or manual execution. An optional disabled schedule trigger (poll_unassigned_incidents) is provided for batch triage sweeps.
  2. Payload Logging (log_incident_payload): Records the incident identifier, impacted service, incoming severity, environment, and error message in the execution trace.
  3. Multi-Tier Triage (triage_incident_with_fallback): The Python script evaluates the error message and diagnostic logs through the OpenAI -> Gemini -> Heuristic failover chain, emitting structured outputs (severity_level, root_cause_hypothesis, blast_radius, mitigation_steps, requires_human_approval, provider_used, failover_occurred).
  4. Slack Broadcast (notify_triage_channel): Posts a formatted Slack message with block elements detailing the diagnosis, blast radius, provider utilized, and recommended mitigation steps.
  5. Approval Gate (evaluate_approval_gate): If severity is HIGH or CRITICAL, or if the triage model flags requires_human_approval: true, the workflow notifies Slack and pauses execution. Responders resume with approval status, approver email, and remediation action.
  6. Mitigation Execution (execute_mitigation_decision): If approved, logs and confirms the mitigation runbook start in Slack. If rejected or timed out, logs that automated remediation stood down.
  7. Audit Record (record_triage_audit): Logs final execution metadata, including the LLM provider utilized and failover state.
  8. Failure Handling (errors): Any unhandled task failure immediately dispatches an urgent failure alert to Slack.

Prerequisites

  • A Slack workspace with an Incoming Webhook configured.
  • An OpenAI API account with an active API key.
  • A Google AI Studio or Google Cloud account with a Gemini API key.
  • Access to Kestra to configure namespace secrets and execute workflows.

Secret Configuration

Configure the following secrets in your Kestra namespace (company.team):

  • OPENAI_API_KEY: API secret key for OpenAI (Tier 1 primary provider).
  • GEMINI_API_KEY: API key for Google Gemini REST API (Tier 2 failover provider).
  • SLACK_WEBHOOK_URL: Slack Incoming Webhook URL for incident triage reports and approval requests.
  • INCIDENT_WEBHOOK_KEY: Secret authentication key used to authorize inbound alert webhooks.

Input Parameters

  • incident_id (STRING, default: "INC-84920"): Unique tracking identifier for the incident or alert.
  • service (STRING, default: "checkout-payment-api"): Impacted service or infrastructure component.
  • severity (SELECT, default: HIGH): Inbound alert priority (CRITICAL, HIGH, MEDIUM, LOW).
  • error_message (STRING): Summary error or threshold breach alert text.
  • logs (STRING): Recent diagnostic logs, stack traces, or metrics context.
  • environment (STRING, default: "production"): Target environment under incident investigation.

Tasks Summary

  • log_incident_payload (io.kestra.plugin.core.log.Log): Logs incoming alert metadata.
  • triage_incident_with_fallback (io.kestra.plugin.scripts.python.Script): Multi-tier LLM triage with OpenAI, Gemini, and offline rule fallback.
  • notify_triage_channel (io.kestra.plugin.notifications.slack.SlackIncomingWebhook): Sends formatted diagnostic summary to Slack.
  • evaluate_approval_gate (io.kestra.plugin.core.flow.If): Conditional branch that gates remediation on human approval:
    • notify_approval_requested (io.kestra.plugin.notifications.slack.SlackIncomingWebhook): Slack alert requesting engineer sign-off.
    • pause_for_human_approval (io.kestra.plugin.core.flow.Pause): Pauses workflow execution for up to 1 hour.
    • execute_mitigation_decision (io.kestra.plugin.core.flow.If): Evaluates resume inputs to trigger remediation or stand down.
  • record_triage_audit (io.kestra.plugin.core.log.Log): Records completion status and provider lineage.
  • alert_workflow_failure (io.kestra.plugin.notifications.slack.SlackIncomingWebhook): Global flow-level error handler.

Example Execution Output

When executed, outputs.triage_incident_with_fallback.vars yields structured outputs:

{
  "incident_id": "INC-84920",
  "service": "checkout-payment-api",
  "provider_used": "OpenAI (gpt-4o-mini)",
  "failover_occurred": false,
  "severity_level": "HIGH",
  "root_cause_hypothesis": "Database connection pool starvation on db-replica-02 causing upstream gateway timeouts on charge endpoints.",
  "blast_radius": "Checkout payment processing degraded for approximately 35% of inbound user checkout requests in us-east-1.",
  "mitigation_steps": "Step 1: Increase connection pool ceiling | Step 2: Recycle checkout worker pods | Step 3: Shed non-critical read replica traffic",
  "summary": "Connection pool exhaustion on db-replica-02 is causing 504 timeouts on POST /api/v2/charge. Responders should verify replica health and recycle worker pods.",
  "requires_human_approval": true
}

Slack notification dispatched:

Incident Triage: INC-84920
Service: checkout-payment-api | Severity: HIGH
LLM Provider: OpenAI (gpt-4o-mini) | Failover Active: False
Root Cause Hypothesis: Database connection pool starvation on db-replica-02...
Blast Radius: Checkout payment processing degraded for approximately 35% of inbound requests...

Quick Start

  1. Add the four required secrets (OPENAI_API_KEY, GEMINI_API_KEY, SLACK_WEBHOOK_URL, INCIDENT_WEBHOOK_KEY) to the company.team namespace in Kestra.
  2. Deploy the blueprint YAML to your Kestra instance.
  3. Trigger a test run manually via the Kestra UI with the default inputs, or send an HTTP POST request to your Kestra webhook endpoint: curl -X POST https://kestra.example.com/api/v1/executions/webhook/company.team/autonomous-incident-triage-with-fallback/{INCIDENT_WEBHOOK_KEY} -H "Content-Type: application/json" -d '{"incident_id": "INC-1001", "service": "auth-service", "severity": "HIGH", "error_message": "Token verification failure spike", "logs": "JWT signature validation timeout"}'
  4. Review the Slack alert in your incident channel.
  5. Navigate to the paused execution in the Kestra UI, click Resume, set approved: true, and confirm mitigation dispatch.

How to Extend

  • Automated Remediation: Replace log_mitigation_approved with a task that triggers an Ansible playbook, Kubernetes pod rollout (io.kestra.plugin.kubernetes.pod.Delete), or AWS ECS service update.
  • PagerDuty Bidirectional Sync: Add a task to acknowledge or update the PagerDuty incident notes with the LLM triage findings.
  • Post-Mortem Storage: Export the incident triage summary and latency metrics to BigQuery or Elasticsearch for weekly incident reviews and SLA tracking.
  • Anthropic Claude Tier: Add Claude 3.5 Sonnet as an intermediate failover tier between OpenAI and Gemini for deep code diff analysis.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.