RealtimeTrigger icon
Switch icon
Get icon
If icon
OutputValues icon
Create icon
Set icon
Invite icon
Post icon
Parallel icon
Create icon
TelegramSend icon
Log icon
Comment icon
Request icon
Delete icon
Pause icon
Archive icon

Open a Slack War Room for Every Google Cloud Monitoring Incident

Open a Slack war room for every Google Cloud Monitoring incident, page the owning team, track it in GitHub, and archive it on resolution.

Categories
CloudInfrastructure

When a Cloud Monitoring alert fires, the first minutes usually go to busywork: creating a channel, finding who is on call, digging out the runbook and opening a ticket, while production is still down. This blueprint does all of that the moment GCP opens an incident, pulls in the team that owns the alert, and cleans up after itself when GCP resolves it.

How it works

  1. Alerting policies send notifications to a Pub/Sub notification channel. The gcp_alert realtime trigger starts one execution per message, on open and on close, and route_by_state (Switch) branches on incident.state.
  2. On open, unless the KV store already holds the incident (so a re-delivered message never opens a second room):
    • route_oncall reads the team from the policy's oncall_label user label, falling back to default.
    • create_war_room and invite_oncall create #inc-<policy>-<incident id> and invite that team.
    • post_briefing posts the severity, summary, observed value against the threshold, resource, who was paged, the runbook from the policy's Documentation field, and a Cloud Console link.
    • track_and_escalate (Parallel) opens and links a GitHub issue (enable_github) and pages Telegram for the severities in escalate_severities (enable_telegram).
    • remember_war_room saves the channel in the KV store as soon as it exists, and remember_tracking_issue adds the issue, so the close can still archive the room if a later step fails.
  3. On close, the flow posts the time to resolve, comments on and closes the issue, waits archive_after in a Pause you can resume early from the UI, then archives the channel.
  4. If any step fails, the errors block alerts on Telegram, or by Slack DM to the first default on-call engineer when Telegram is disabled.

What you get

  • A channel per incident with the owning team inside and the runbook posted. Label a policy team = payments and its incidents page the payments team.
  • Optional GitHub tracking and severity-aware Telegram paging. Set enable_github and enable_telegram to false for a Slack-only setup.
  • Clean-up and idempotency: resolved incidents close their issue and archive their channel, and duplicate notifications are ignored.
  • Flow outputs, so a downstream flow (for example a postmortem drafter on a Flow trigger) knows which room and issue belong to the incident:
    • incident_id and incident_state: the Cloud Monitoring incident, and whether this run handled its opening or its closing.
    • slack_channel_id: the war room channel.
    • github_issue_number: the tracking issue, or 0 when GitHub is disabled.

Who it's for

SRE and platform teams on Google Cloud that already alert with Cloud Monitoring, especially when several teams share on-call and each alert should reach the team that owns it.

Why orchestrate this with Kestra

An incident is a lifecycle, not a single notification. Kestra reacts to Cloud Monitoring in real time through the Pub/Sub trigger, correlates the open and close events of the same incident through its KV store, holds the room open in a Pause anyone can resume from the UI, and alerts through its errors block when the automation itself breaks. A Cloud Function or a cron job can post a message, but following an incident from open to archive would also need its own state store, retries and failure alerting.

Prerequisites

PROJECT_ID=your-project
PROJECT_NUMBER=$(gcloud projects describe $PROJECT_ID --format='value(projectNumber)')
SA=kestra-incident@$PROJECT_ID.iam.gserviceaccount.com

# Topic for the alerting policies, and the subscription Kestra reads
gcloud pubsub topics create incident-alerts --project $PROJECT_ID
gcloud pubsub subscriptions create incident-alerts-kestra --topic incident-alerts --project $PROJECT_ID

# Pub/Sub notification channel (add it to your alerting policies), allowed to publish to the topic
gcloud beta monitoring channels create --type pubsub --display-name "Kestra incident war room" \
  --channel-labels topic=projects/$PROJECT_ID/topics/incident-alerts --project $PROJECT_ID
gcloud pubsub topics add-iam-policy-binding incident-alerts --project $PROJECT_ID --role roles/pubsub.publisher \
  --member serviceAccount:service-$PROJECT_NUMBER@gcp-sa-monitoring-notification.iam.gserviceaccount.com

# Service account for Kestra, limited to this one subscription; the key becomes GCP_SERVICE_ACCOUNT
gcloud iam service-accounts create kestra-incident --project $PROJECT_ID
gcloud pubsub subscriptions add-iam-policy-binding incident-alerts-kestra --project $PROJECT_ID \
  --member serviceAccount:$SA --role roles/pubsub.subscriber
gcloud iam service-accounts keys create kestra-incident.json --iam-account $SA --project $PROJECT_ID

autoCreateSubscription is off because the subscription already exists. With it on, the trigger lists every subscription in the project at startup, which needs project-wide Pub/Sub permissions this service account deliberately doesn't have.

On each alerting policy, write the runbook in Documentation and add a user label for the owning team, for example team = payments.

  • Slack: an app with the bot scopes channels:manage, channels:read and chat:write. Copy its xoxb-… token.
  • GitHub (optional): a fine-grained token limited to the tracking repository, with Issues: Read and write.
  • Telegram (optional): a bot from @BotFather and the on-call chat ID. Send the bot a message first: bots cannot start conversations.

Secrets

  • GCP_PROJECT_ID: project that hosts the topic and subscription.
  • GCP_SERVICE_ACCOUNT: JSON key of the kestra-incident service account.
  • SLACK_TOKEN: Slack bot token.
  • GITHUB_TOKEN: GitHub token, only with enable_github.
  • TELEGRAM_TOKEN and TELEGRAM_CHAT_ID: Telegram bot token and chat, only with enable_telegram.

Variables

  • pubsub_topic, pubsub_subscription: the topic and subscription created above.
  • slack_oncall_by_team: team name to Slack member IDs (Profile → ⋮ → Copy member ID). A default entry with at least one ID is required.
  • oncall_label: the alerting policy user label that holds the team. Defaults to team.
  • enable_github, github_repository: open a tracking issue per incident in owner/repo.
  • enable_telegram, escalate_severities: page on Telegram for these severities, written in lowercase. Defaults to critical and error.
  • archive_after: ISO 8601 duration a resolved room stays open. Defaults to PT24H.

Quick start

  1. Complete the prerequisites, add the secrets to your namespace and set the variables.

  2. Save the flow; the realtime trigger starts listening immediately.

  3. Test without breaking anything by publishing a sample incident:

    gcloud pubsub topics publish incident-alerts --project $PROJECT_ID --message '{"version":"1.2","incident":{"incident_id":"0.demo123","state":"open","started_at":1767225600,"policy_name":"Demo policy","condition_name":"Demo condition","summary":"Demo incident","severity":"Critical","url":"https://console.cloud.google.com/monitoring/alerting","documentation":{"content":"Demo runbook"}}}'
    

    Publish it again with "state":"closed" and an ended_at timestamp to run the resolution branch, and add "policy_user_labels":{"team":"payments"} to try team routing.

How to extend

  • Look up the current on-call from PagerDuty or Opsgenie in route_oncall instead of the slack_oncall_by_team map.
  • Attach the last hour of the alerting metric to the briefing with io.kestra.plugin.gcp.monitoring.Query.
  • Chain a postmortem drafter on the flow outputs with an io.kestra.plugin.core.trigger.Flow trigger.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.