New to Kestra?
Use blueprints to kickstart your first workflows.
Deduplicate Kestra workflow failure alerts using KV store state and route incident notifications to Matrix rooms by namespace severity.
When automated data pipelines or microservice workflows fail across multiple environments, operational teams face alerting noise and delayed response times. Sending an instant message for every repeated failure leads to alert fatigue, while sending all alerts to a single chat channel obscures critical production outages among routine development warnings.
This operational blueprint provides a centralized Matrix incident dispatcher for Kestra. It monitors workflow executions, deduplicates repeated alerts using Kestra KV store state, and dynamically routes failure notifications to environment-specific Matrix rooms (#prod-incidents, #dev-alerts, #general-alerts).
on_failure trigger (io.kestra.plugin.core.trigger.Flow) listens for FAILED or WARNING execution states across all company.* namespaces. The condition trigger.flowId != flow.id prevents self-triggering loops.check_recent_alert (io.kestra.plugin.core.kv.Get) looks up the sanitized key dedup_{namespace}_{flowId} in Kestra KV store with errorOnMissing: false.evaluate_alert (io.kestra.plugin.core.flow.If) evaluates outputs.check_recent_alert.value == null:route_by_namespace (io.kestra.plugin.core.flow.Switch) performs exact string matching on trigger.namespace (e.g. company.prod, company.dev). Any namespace not explicitly matched falls back to defaults and dispatches to default_room_id.NOTICE message to the designated Matrix room using io.kestra.plugin.matrix.Send and secret MATRIX_ACCESS_TOKEN.record_alert_cooldown (io.kestra.plugin.core.kv.Set) writes the deduplication key with a TTL duration specified by inputs.cooldown (default PT15M).log_suppressed_alert (io.kestra.plugin.core.log.Log) logs a suppression notice and skips sending a chat message.errors block logs dispatcher runtime issues without calling external chat endpoints to avoid circular notification loops.Centralizing failure dispatching inside Kestra eliminates redundant webhook configurations in individual workflows. Using a Flow trigger combined with KV state deduplication ensures that temporary workflow retry loops or cascading job failures do not spam chat channels.
company.staging, company.data) to the Switch task to route to dedicated team rooms.cooldown duration input per environment to customize alert suppression windows.https://matrix.org or a self-hosted Synapse server).inputs defaults.MATRIX_ACCESS_TOKEN: Access token for the Matrix bot account.Add the secret MATRIX_ACCESS_TOKEN to your Kestra namespace.
Create an unencrypted room in Matrix (e.g., #prod-incidents), invite your bot, and copy the internal room ID (e.g., !abc123xyz:matrix.org from Room Settings > Advanced).
Import this blueprint and set prod_room_id, dev_room_id, and default_room_id input defaults.
Test the dispatcher by executing this sample failing flow in a watched namespace:
id: sample-failing-job
namespace: company.prod
tasks:
- id: trigger_failure
type: io.kestra.plugin.core.execution.Fail
message: "Simulated production pipeline failure for Matrix alert test."
Execute sample-failing-job. The dispatcher will catch the failure, post a notice to Matrix, and set a 15-minute KV deduplication lock.
check_recent_alert returns null on first run.NOTICE message with direct execution link ({{ kestra.url }}/ui/executions/...).dedup_company_prod_sample-failing-job with PT15M expiration.sample-failing-job within 15 minutes logs "Alert for flow sample-failing-job in namespace company.prod was suppressed".trigger.* context variables (trigger.namespace, trigger.flowId, trigger.executionId, trigger.state) are only populated during event-driven trigger executions. Always test by executing a sample failing flow.M_FORBIDDEN.!id:domain), not room aliases starting with #.