Query icon
If icon
ChatCompletion icon
GoogleGemini icon
SlackIncomingWebhook icon
Log icon
RealtimeTrigger icon

Manufacturing Downtime Root-Cause Triage with AI and MQTT

Monitor manufacturing downtime from MQTT events, query machine stop history in PostgreSQL, classify long stops with Gemini AI, and alert operators in Slack.

Categories
AIInfrastructure

This blueprint automates the root-cause analysis of machine downtime in manufacturing environments. When a production line stops (triggered by an MQTT event from a PLC or sensor), the workflow:

  1. Captures the event via MQTT RealtimeTrigger (e.g., plant/filler1/stops).
  2. Queries historical context from Postgres: recent stops on the same machine.
  3. Applies conditional logic: Micro-stops (< 2 min) are logged; longer stops trigger AI analysis.
  4. Analyzes root causes using Google Gemini to synthesize event data + history into a likely cause, suggested check, and severity.
  5. Stores the summary back to the downtime_log table for audit trails.
  6. Alerts the team via Slack with the AI-generated summary, enabling faster response.

How it works

Architecture

  • Trigger: MQTT RealtimeTrigger listens on plant/+/stops for machine stop events (JSON payloads).
  • Task 1 (ensure_table): Idempotent table creation in Postgres for durability.
  • Task 2 (fetch_history): Queries recent downtime records for the affected machine to provide context.
  • Task 3 (log_event): Inserts the current stop into the audit log immediately.
  • Task 4 (is_long_stop): Flowable If condition, branching on stop duration.
    • Then branch: For stops of at least 2 minutes, runs AI analysis and sends a Slack alert.
    • Else branch: For micro-stops, logs a message without running AI.

Sample MQTT Payload

{
  "machine_id": "filler1",
  "line": "CSD3",
  "reason_code": "LOW_PRESSURE",
  "duration_min": 7
}

AI Prompt

The Gemini ChatCompletion receives the current stop, the machine's recent stop history, and generates:

  • Likely cause (e.g., "Recurring pressure regulation issue")
  • Suggested check (e.g., "Inspect pressure relief valve on filler1")
  • Severity (LOW, MEDIUM, HIGH)

Slack Alert Format

Includes machine ID, line, duration, reason code, and the AI-generated summary, formatted as a rich Block Kit message.

What you get

  • Event-driven monitoring: No polling, no cron. Reacts instantly to MQTT events.
  • Audit trail: Every downtime event persists to Postgres with AI analysis attached.
  • Contextual AI: Gemini analyzes patterns, not isolated events.
  • Conditional logic: Reduces noise by ignoring micro-stops.
  • Slack integration: Team gets instant, actionable alerts.
  • Extensible: Easy to add email, PagerDuty, or Kafka notifications.

Who it's for

  • Manufacturing plants running PLCs and IoT sensors.
  • Reliability engineers who need faster root-cause analysis cycles.
  • Operations teams who want visibility into downtime patterns.
  • Anyone integrating MQTT-based machinery with a data warehouse.

Prerequisites

  • MQTT broker (e.g., Mosquitto) publishing machine stop events.
  • Postgres database for historical downtime logs.
  • Google Gemini API key for AI analysis.
  • Slack workspace with an incoming webhook URL for alerts.
  • Kestra instance with MQTT, JDBC PostgreSQL, and Slack notification plugins.

Environment setup

Use the included docker-compose.yml to spin up a local MQTT broker and Postgres:

cd kestra-dev/
docker compose up -d

This starts:

  • Mosquitto MQTT broker on localhost:1883
  • PostgreSQL on localhost:5433 (user: plant, password: plant, db: plant)

Secrets

In the Kestra UI, add the following namespace secrets (or set as environment variables prefixed SECRET_):

Secret Example Source
PLANT_DB_USER plant Postgres username
PLANT_DB_PASSWORD plant Postgres password
GEMINI_API_KEY AIza... Google AI Studio
SLACK_WEBHOOK_URL https://hooks.slack.com/... Slack App incoming webhook

Quick start

  1. Deploy MQTT + Postgres:

    cd kestra-dev/
    docker compose up -d
    
  2. Add secrets to Kestra (Namespace → Secrets):

    • PLANT_DB_USER = plant
    • PLANT_DB_PASSWORD = plant
    • GEMINI_API_KEY = your API key
    • SLACK_WEBHOOK_URL = your Slack webhook
  3. Copy both flows (downtime-root-cause-triage.yaml and simulate-machine-stops.yaml) into the Kestra editor under the company.team namespace.

  4. Trigger a test:

    • Manually execute simulate-machine-stops, which publishes two MQTT events (a 7-min stop and a 1-min stop).
    • Watch the Kestra execution topology; you should see:
      • is_long_stop branches on the 7-minute stop (goes to then, runs AI).
      • is_long_stop evaluates to false on the 1-minute stop (goes to else, just logs).
    • Check Postgres for 2 new rows in downtime_log.
    • Check Slack for 1 alert (only the long stop triggers an alert).
  5. Inspect outputs:

    • Click the analyze task to see the AI-generated summary.
    • Click alert to confirm the Slack message was sent.
    • Query Postgres directly to see the full audit log.

Scaling and customization

  • Adjust long_stop_minutes (default 2) to change the cutoff for AI analysis.
  • Increase history_limit (default 10) to give Gemini more historical context.
  • Extend the payload with additional fields (e.g., OEE, shift ID, production target) and reference them in the AI prompt.
  • Integrate with your PLC: Update the MQTT topic and payload schema to match your sensors.

Tags

  • Manufacturing
  • IoT
  • MQTT
  • AI
  • Postgres
  • Slack
  • Event-Driven
  • Reliability Engineering
See How

New to Kestra?

Use blueprints to kickstart your first workflows.