This blueprint automates the root-cause analysis of machine downtime in manufacturing environments.
When a production line stops (triggered by an MQTT event from a PLC or sensor), the workflow:
- Captures the event via MQTT RealtimeTrigger (e.g.,
plant/filler1/stops).
- Queries historical context from Postgres: recent stops on the same machine.
- Applies conditional logic: Micro-stops (< 2 min) are logged; longer stops trigger AI analysis.
- Analyzes root causes using Google Gemini to synthesize event data + history into a likely cause, suggested check, and severity.
- Stores the summary back to the downtime_log table for audit trails.
- Alerts the team via Slack with the AI-generated summary, enabling faster response.
How it works
Architecture
- Trigger: MQTT RealtimeTrigger listens on
plant/+/stops for machine stop events (JSON payloads).
- Task 1 (ensure_table): Idempotent table creation in Postgres for durability.
- Task 2 (fetch_history): Queries recent downtime records for the affected machine to provide context.
- Task 3 (log_event): Inserts the current stop into the audit log immediately.
- Task 4 (is_long_stop): Flowable If condition, branching on stop duration.
- Then branch: For stops of at least 2 minutes, runs AI analysis and sends a Slack alert.
- Else branch: For micro-stops, logs a message without running AI.
Sample MQTT Payload
{
"machine_id": "filler1",
"line": "CSD3",
"reason_code": "LOW_PRESSURE",
"duration_min": 7
}
AI Prompt
The Gemini ChatCompletion receives the current stop, the machine's recent stop history, and generates:
- Likely cause (e.g., "Recurring pressure regulation issue")
- Suggested check (e.g., "Inspect pressure relief valve on filler1")
- Severity (LOW, MEDIUM, HIGH)
Slack Alert Format
Includes machine ID, line, duration, reason code, and the AI-generated summary, formatted as a rich Block Kit message.
What you get
- Event-driven monitoring: No polling, no cron. Reacts instantly to MQTT events.
- Audit trail: Every downtime event persists to Postgres with AI analysis attached.
- Contextual AI: Gemini analyzes patterns, not isolated events.
- Conditional logic: Reduces noise by ignoring micro-stops.
- Slack integration: Team gets instant, actionable alerts.
- Extensible: Easy to add email, PagerDuty, or Kafka notifications.
Who it's for
- Manufacturing plants running PLCs and IoT sensors.
- Reliability engineers who need faster root-cause analysis cycles.
- Operations teams who want visibility into downtime patterns.
- Anyone integrating MQTT-based machinery with a data warehouse.
Prerequisites
- MQTT broker (e.g., Mosquitto) publishing machine stop events.
- Postgres database for historical downtime logs.
- Google Gemini API key for AI analysis.
- Slack workspace with an incoming webhook URL for alerts.
- Kestra instance with MQTT, JDBC PostgreSQL, and Slack notification plugins.
Environment setup
Use the included docker-compose.yml to spin up a local MQTT broker and Postgres:
cd kestra-dev/
docker compose up -d
This starts:
- Mosquitto MQTT broker on
localhost:1883
- PostgreSQL on
localhost:5433 (user: plant, password: plant, db: plant)
Secrets
In the Kestra UI, add the following namespace secrets (or set as environment variables prefixed SECRET_):
| Secret |
Example |
Source |
PLANT_DB_USER |
plant |
Postgres username |
PLANT_DB_PASSWORD |
plant |
Postgres password |
GEMINI_API_KEY |
AIza... |
Google AI Studio |
SLACK_WEBHOOK_URL |
https://hooks.slack.com/... |
Slack App incoming webhook |
Quick start
Deploy MQTT + Postgres:
cd kestra-dev/
docker compose up -d
Add secrets to Kestra (Namespace → Secrets):
PLANT_DB_USER = plant
PLANT_DB_PASSWORD = plant
GEMINI_API_KEY = your API key
SLACK_WEBHOOK_URL = your Slack webhook
Copy both flows (downtime-root-cause-triage.yaml and simulate-machine-stops.yaml) into the Kestra editor under the company.team namespace.
Trigger a test:
- Manually execute
simulate-machine-stops, which publishes two MQTT events (a 7-min stop and a 1-min stop).
- Watch the Kestra execution topology; you should see:
is_long_stop branches on the 7-minute stop (goes to then, runs AI).
is_long_stop evaluates to false on the 1-minute stop (goes to else, just logs).
- Check Postgres for 2 new rows in
downtime_log.
- Check Slack for 1 alert (only the long stop triggers an alert).
Inspect outputs:
- Click the
analyze task to see the AI-generated summary.
- Click
alert to confirm the Slack message was sent.
- Query Postgres directly to see the full audit log.
Scaling and customization
- Adjust
long_stop_minutes (default 2) to change the cutoff for AI analysis.
- Increase
history_limit (default 10) to give Gemini more historical context.
- Extend the payload with additional fields (e.g., OEE, shift ID, production target) and reference them in the AI prompt.
- Integrate with your PLC: Update the MQTT topic and payload schema to match your sensors.
Tags
Manufacturing
IoT
MQTT
AI
Postgres
Slack
Event-Driven
Reliability Engineering