New to Kestra?
Use blueprints to kickstart your first workflows.
Monitors Kubernetes cluster node conditions to proactively detect DiskPressure, MemoryPressure, and PIDPressure before pod evictions cascade.
This blueprint acts as a proactive reliability sentinel across Kubernetes clusters. Scheduled every 15 minutes, it queries the Kubernetes API to evaluate node condition flags (MemoryPressure, DiskPressure, PIDPressure, and NetworkUnavailable), alerting platform SRE teams to degrading nodes before the kubelet triggers disruptive workload evictions or enters a NotReady state.
periodic_k8s_node_probe trigger (io.kestra.plugin.core.trigger.Schedule) initiates node inspections every 15 minutes.inspect_node_conditions task (io.kestra.plugin.scripts.python.Script) connects to the cluster using credentials provided in KUBE_CONFIG_DATA, querying the CoreV1Api to parse each node's status conditions.check_pressure_conditions flowable task (io.kestra.plugin.core.flow.If) branches based on whether any nodes are actively flagging resource pressures.notify_slack_sre (io.kestra.plugin.notifications.slack.SlackIncomingWebhook) dispatches an urgent incident notification with hostnames.log_nodes_healthy records compliant cluster capacity.export_node_manifest task annotates the execution with searchable status labels.flowchart TD
A[Schedule: Every 15 Minutes] --> B[inspect_node_conditions: Kubernetes CoreV1Api]
B --> C{Node Pressure Active?}
C -- Yes --> D[notify_slack_sre: Slack Webhook]
C -- No --> E[log_nodes_healthy: Log]
D --> F[export_node_manifest: Execution Labels]
E --> F
| Name | Type | Default | Description |
|---|---|---|---|
cluster_name |
STRING | production-eks-us-east-1 |
Identifier name of the Kubernetes cluster being monitored. |
alert_on_any_pressure |
BOOL | true |
Whether to trigger an immediate Slack alert when any node reports resource pressure conditions. |
slack_channel |
STRING | #k8s-platform-alerts |
Slack channel destination for Kubernetes cluster node health alerts. |
{{ outputs.inspect_node_conditions.vars.total_nodes_inspected }}: Total nodes evaluated.{{ outputs.inspect_node_conditions.vars.degraded_nodes_count }}: Number of nodes flagging resource pressure.{{ outputs.inspect_node_conditions.vars.has_node_pressure }}: Boolean flag indicating cluster degradation.{{ outputs.inspect_node_conditions.vars.degraded_node_names }}: List of affected node names.{{ outputs.inspect_node_conditions.outputFiles['k8s_node_pressure_report.json'] }}: Full JSON node condition diagnostics.kubeconfig with read permissions on nodes (get, list).KUBE_CONFIG_DATA: Base64 or plain YAML kubeconfig configuration string.SLACK_WEBHOOK_URL: Slack Incoming Webhook endpoint URL.KUBE_CONFIG_DATA and SLACK_WEBHOOK_URL in your Kestra namespace secrets.cluster_name.NetworkUnavailable; ensure nodes remain under observation before initiating manual cordoning.current-context set properly if containing multi-cluster definitions.kubectl cordon <node> to prevent the scheduler from placing new pods on under-pressure nodes.DiskPressure is flagged.