Schedule icon
Request icon
If icon
SlackIncomingWebhook icon
Log icon
Return icon

Elasticsearch Cluster Health and Shard Allocation Sentinel

Automated Elasticsearch monitoring workflow to query _cluster/health, detect yellow/red status and unassigned shards, and alert Slack.

Categories
BusinessData

In Elasticsearch deployments supporting log analytics, full-text search, and observability pipelines, maintaining a green cluster state is essential for query availability and data durability. When nodes drop offline, disk high-watermarks are breached, or indices fail allocation, Elasticsearch enters yellow (missing replica shards) or red (missing primary shards) states.

Left undetected, unassigned shards cause write rejections, degraded search latency, and potential data loss if an underlying node hardware failure occurs while replicas remain unallocated.

This blueprint establishes an automated health watchdog for Elasticsearch. Executing every 10 minutes (or on-demand), it queries the native /_cluster/health API, inspects overall cluster status, monitors unassigned and relocating shard counts, and immediately delivers diagnostic cards to Slack when degradation is detected.

How it works

  1. Scheduled Cluster Audit: The scheduled_cluster_audit trigger (io.kestra.plugin.core.trigger.Schedule) executes every 10 minutes or on-demand via the Kestra UI.
  2. Health Request: The check_cluster_health task (io.kestra.plugin.elasticsearch.Request) calls GET /_cluster/health against your target Elasticsearch endpoint with basic authentication.
  3. Status Evaluation: The evaluate_cluster_status flowable task (io.kestra.plugin.core.flow.If) branches based on whether the cluster status is non-green or unassigned shards exceed max_allowed_unassigned_shards.
  4. Slack Alert Dispatch: If degradation is present, notify_sre_slack (io.kestra.plugin.slack.notifications.SlackIncomingWebhook) broadcasts an alert with cluster name, active shard percentage, unassigned count, and troubleshooting guidance.
  5. Healthy Confirmation: When cluster status is green, log_healthy_cluster records nominal metrics in execution logs.
  6. Audit Manifest: The export_health_manifest task records telemetry outputs for infrastructure uptime tracking.

What you get

  • Automated detection of unassigned shards and cluster state degradation.
  • Early warning before missing replicas turn into catastrophic primary shard data loss.
  • Visibility into active shard percentages and node count shifts.
  • Immediate remediation hints pointing to the Elasticsearch allocation explain API.

Who it is for

  • Site Reliability Engineers and DevOps teams maintaining Elasticsearch or OpenSearch clusters.
  • Platform engineers operating ELK log aggregation and distributed tracing infrastructure.
  • Database Administrators responsible for search index durability and high availability.

Why orchestrate this with Kestra

Monitoring Elasticsearch health usually requires complex metric daemons or heavy external monitoring agents. Kestra provides declarative, lightweight orchestration: it executes authenticated HTTP requests against native REST APIs, schedules continuous sweeps, conditionally alerts the team on Slack, and securely handles credentials.

Inputs

Name Type Default Description
elasticsearch_host STRING http://elasticsearch:9200 HTTP endpoint of the target Elasticsearch cluster.
max_allowed_unassigned_shards INT 0 Tolerance threshold for unassigned shards before triggering an alert.
slack_channel STRING #sre-alerts Slack channel destination for Elasticsearch cluster alerts.

Expected outputs

  • {{ outputs.check_cluster_health.response.status }}: Overall cluster health status (green, yellow, or red).
  • {{ outputs.check_cluster_health.response.unassigned_shards }}: Total number of primary or replica shards not currently allocated to nodes.
  • {{ outputs.check_cluster_health.response.active_shards_percent_as_number }}: Percentage of active shards currently online.
  • {{ outputs.export_health_manifest.value }}: Structured JSON telemetry manifest recording execution timestamp and cluster metrics.

Prerequisites

  • An Elasticsearch (version 7.x or 8.x) or OpenSearch cluster.
  • A service account user with monitor cluster privileges to query /_cluster/health.
  • A Slack Incoming Webhook URL targeting your SRE notification channel.

Secrets

  • ELASTICSEARCH_USER: Username of the account authorized to read cluster health.
  • ELASTICSEARCH_PASSWORD: Password for the Elasticsearch monitoring user.
  • SLACK_WEBHOOK_URL: Slack Incoming Webhook endpoint URL used for sending alert notifications.

Quick start

  1. Configure ELASTICSEARCH_USER, ELASTICSEARCH_PASSWORD, and SLACK_WEBHOOK_URL in your Kestra namespace secrets.
  2. Import this flow YAML into your Kestra instance.
  3. Click Execute in the UI to run an on-demand cluster health probe.
  4. Review execution outputs to check cluster status and shard allocation ratios.

Common pitfalls and troubleshooting

  • Yellow Status on Single-Node Clusters: A single-node cluster will naturally remain in yellow status if indices are configured with number_of_replicas: 1 because replicas cannot be allocated to the same node. Set inputs.max_allowed_unassigned_shards accordingly for development environments.
  • Disk Watermark Lockouts: Elasticsearch blocks shard allocation when disk usage exceeds the high watermark (default 90%). Check disk capacity on data nodes if shards remain unassigned.
  • HTTPS & Self-Signed Certificates: If your cluster enforces TLS with internal certificates, ensure the host URL matches your CA configuration.

How to extend

  • Add an automated diagnostic task calling GET /_cluster/allocation/explain to embed root-cause explanations directly into the Slack message.
  • Integrate with PagerDuty for critical red cluster status notifications.
  • Add index lifecycle management checks to identify indices lacking ILM policies.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.