Script icon
If icon
SlackIncomingWebhook icon
Log icon
Labels icon
Schedule icon

Kubernetes Node Resource Pressure Sentinel

Monitors Kubernetes cluster node conditions to proactively detect DiskPressure, MemoryPressure, and PIDPressure before pod evictions cascade.

Categories
CoreInfrastructure

This blueprint acts as a proactive reliability sentinel across Kubernetes clusters. Scheduled every 15 minutes, it queries the Kubernetes API to evaluate node condition flags (MemoryPressure, DiskPressure, PIDPressure, and NetworkUnavailable), alerting platform SRE teams to degrading nodes before the kubelet triggers disruptive workload evictions or enters a NotReady state.

How it works

  1. Recurring 15-Minute Probe: The periodic_k8s_node_probe trigger (io.kestra.plugin.core.trigger.Schedule) initiates node inspections every 15 minutes.
  2. Kubernetes API Node Inspection: The inspect_node_conditions task (io.kestra.plugin.scripts.python.Script) connects to the cluster using credentials provided in KUBE_CONFIG_DATA, querying the CoreV1Api to parse each node's status conditions.
  3. Condition Gate: The check_pressure_conditions flowable task (io.kestra.plugin.core.flow.If) branches based on whether any nodes are actively flagging resource pressures.
  4. Slack SRE Dispatch: When degraded nodes are detected, notify_slack_sre (io.kestra.plugin.notifications.slack.SlackIncomingWebhook) dispatches an urgent incident notification with hostnames.
  5. Healthy Confirmation: When all nodes report healthy status, log_nodes_healthy records compliant cluster capacity.
  6. Manifest Export: The export_node_manifest task annotates the execution with searchable status labels.

Architecture diagram

flowchart TD
    A[Schedule: Every 15 Minutes] --> B[inspect_node_conditions: Kubernetes CoreV1Api]
    B --> C{Node Pressure Active?}
    C -- Yes --> D[notify_slack_sre: Slack Webhook]
    C -- No --> E[log_nodes_healthy: Log]
    D --> F[export_node_manifest: Execution Labels]
    E --> F

Use cases

  • Proactive Pod Eviction Prevention: Address container runtime disk bloat or memory leaks before the kubelet evicts production databases or queues.
  • PID Starvation Detection: Catch fork-bombing workloads or thread pool leaks before node PID limits are exhausted.
  • Fleet-Wide Cluster Health Auditing: Run automated health probes across distributed multi-region Kubernetes clusters.

Inputs

Name Type Default Description
cluster_name STRING production-eks-us-east-1 Identifier name of the Kubernetes cluster being monitored.
alert_on_any_pressure BOOL true Whether to trigger an immediate Slack alert when any node reports resource pressure conditions.
slack_channel STRING #k8s-platform-alerts Slack channel destination for Kubernetes cluster node health alerts.

Expected outputs

  • {{ outputs.inspect_node_conditions.vars.total_nodes_inspected }}: Total nodes evaluated.
  • {{ outputs.inspect_node_conditions.vars.degraded_nodes_count }}: Number of nodes flagging resource pressure.
  • {{ outputs.inspect_node_conditions.vars.has_node_pressure }}: Boolean flag indicating cluster degradation.
  • {{ outputs.inspect_node_conditions.vars.degraded_node_names }}: List of affected node names.
  • {{ outputs.inspect_node_conditions.outputFiles['k8s_node_pressure_report.json'] }}: Full JSON node condition diagnostics.

Prerequisites

  • Kubernetes cluster kubeconfig with read permissions on nodes (get, list).
  • Incoming Webhook URL configured for your platform operations Slack channel.

Secrets

  • KUBE_CONFIG_DATA: Base64 or plain YAML kubeconfig configuration string.
  • SLACK_WEBHOOK_URL: Slack Incoming Webhook endpoint URL.

Quick start

  1. Configure KUBE_CONFIG_DATA and SLACK_WEBHOOK_URL in your Kestra namespace secrets.
  2. Import this flow YAML into your Kestra workspace.
  3. Specify your cluster identifier in cluster_name.
  4. Click Execute in the UI to perform an initial node resource pressure probe.

Common pitfalls and troubleshooting

  • Transient Network Pressure: Transient CNI flakiness can briefly trigger NetworkUnavailable; ensure nodes remain under observation before initiating manual cordoning.
  • Kubeconfig Context: Ensure your kubeconfig has current-context set properly if containing multi-cluster definitions.

How to extend

  • Add automated node cordoning via kubectl cordon <node> to prevent the scheduler from placing new pods on under-pressure nodes.
  • Trigger automated container image cache pruning when DiskPressure is flagged.
  • Route critical memory pressure alerts to PagerDuty or Opsgenie for high-priority response.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.