Webhook icon
Template icon
Upgrade icon
If icon
SlackIncomingWebhook icon
Pause icon
Sequential icon
Status icon
Fail icon
Rollback icon

Gated Helm Release with Automatic Rollback

Deploy a Helm chart with Kestra through a server-side dry run and a production approval gate, with automatic rollback and Slack alerts on failure.

Categories
Infrastructure

Ship Helm releases through the same guarded path every time. This blueprint renders the chart offline, validates the exact upgrade against the live API server, pauses production deploys for a human approval, then runs helm upgrade --install and waits for every resource to become ready. If the deploy, the read-back or the health check fails, the release is rolled back to its previous revision and Slack is alerted. CI triggers it with a webhook, so a pipeline can push a new image tag without anyone touching kubectl or a kubeconfig.

How it works

  1. render_chart (io.kestra.plugin.helm.Template) runs helm template for the podinfo chart 6.15.0 with the requested image tag. It never contacts the cluster. The rendered manifest is stored with the execution ({{ outputs.render_chart.manifest }}), so every run keeps a reviewable, diffable record of what was about to ship.
  2. validate_on_cluster (io.kestra.plugin.helm.Upgrade with dryRun: SERVER) runs the same upgrade as a server-side dry run against the API server. Bad credentials, an unreachable cluster or a chart that will not install stop the flow here, before approval is requested and before anything is applied. A dry run emits no Assets.
  3. production_gate (io.kestra.plugin.core.flow.If) checks inputs.environment. For production, request_approval posts to Slack and approve_production (io.kestra.plugin.core.flow.Pause) holds the run for up to one hour. Resume the execution to deploy. With behavior: FAIL, an approval nobody acts on fails the run instead of deploying. Staging skips the gate.
  4. release (io.kestra.plugin.core.flow.Sequential) does the actual change:
    • deploy (io.kestra.plugin.helm.Upgrade) runs helm upgrade --install with createNamespace: true, wait: WATCHER and timeout: PT10M, so the task only succeeds once the new pods are ready. It registers the release and its resources as Assets tagged with the environment.
    • verify_release (io.kestra.plugin.helm.Status) reads the release back from the cluster.
    • fail_if_unhealthy (io.kestra.plugin.core.execution.Fail) fails the step unless {{ outputs.verify_release.status }} is deployed.
  5. The errors branch on release only runs when one of those three tasks fails. A failed dry run or an expired approval never triggers it. rollback (io.kestra.plugin.helm.Rollback, wait: WATCHER) reverts to the previous revision and waits for it to be ready, then alert_rollback posts the outcome to Slack. The execution still ends FAILED, so the failure stays visible.
  6. notify_success posts the release name, deployed revision and environment to Slack.
  7. The ci_deploy_webhook trigger (io.kestra.plugin.core.trigger.Webhook) lets CI start a deploy. The image tag is read with {{ trigger.body.image_tag ?? inputs.image_tag }}: the JSON body wins when it carries image_tag, and manual runs fall back to the input.

What you get

  • A rendered manifest stored with every execution, for audit and diffing between releases.
  • A server-side dry run that catches cluster-side problems before a human is asked to approve anything.
  • A one-hour human approval gate for production, with a Slack nudge, that fails closed.
  • Readiness-gated deploys: the task waits for pods to be ready, not just for the API server to accept the objects.
  • An automatic rollback scoped to the release step, with a Slack alert reporting the revision it landed on.
  • Helm releases and the Kubernetes resources they manage registered as Assets, with environment metadata.

Who it's for

  • Platform and DevOps teams who deploy Helm charts from CI and want an approval step for production.
  • SREs who want failed releases reverted automatically, with an alert and an audit trail.
  • Teams replacing helm upgrade --atomic shell steps in CI with something they can observe, resume and replay.

Why orchestrate this with Kestra

Helm on its own has no approval step, no notifications, and no memory of who deployed what. A CI job that shells out to helm upgrade loses the rendered manifest when the runner is recycled and cannot wait an hour for a human without holding a runner. Kestra keeps each stage as a separate, retryable task with its own logs and outputs. The approval pause costs nothing while it waits, the rollback is scoped so it only runs for the release step, and every run records the manifest, revision and release Assets. The same flow serves CI (webhook) and humans (manual run with inputs).

Prerequisites

  • A Kubernetes cluster whose API server is reachable from the Kestra worker.
  • A service account token allowed to manage the release namespaces (create namespaces, Deployments, Services, HorizontalPodAutoscalers and the Secrets Helm uses to store release state).
  • A Kestra worker with Docker available. The Helm tasks run the Helm 4 CLI in the alpine/helm:4.3.0 container by default. A Helm 3 image is not supported.
  • The Helm plugin (io.kestra.plugin.helm) and the Slack plugin installed.
  • A Slack incoming webhook URL for the release channel.

Secrets

  • K8S_MASTER_URL: Kubernetes API server URL, e.g. https://my-cluster.example.com:6443.
  • K8S_CA_CERT: the cluster CA certificate, base64-encoded (the certificate-authority-data value from a kubeconfig). Do not paste a raw PEM.
  • K8S_TOKEN: bearer token of the service account Helm deploys with.
  • SLACK_WEBHOOK_URL: Slack incoming webhook used for the approval request, the success message and the rollback alert.
  • HELM_DEPLOY_WEBHOOK_KEY: the secret key in the webhook URL that CI calls. Treat it like a password.

Inputs

  • environment (SELECT, staging or production, default staging): selects the namespace podinfo-<environment> and whether the approval gate applies. Webhook runs use the default (staging).
  • image_tag (STRING, default 6.15.0): podinfo image tag for manual runs. A webhook body field image_tag overrides it.

Outputs

  • {{ outputs.render_chart.manifest }}: URI of the offline-rendered manifest in internal storage.
  • {{ outputs.render_chart.resources }}: kinds, names and namespaces of the resources the chart renders.
  • {{ outputs.deploy.releaseName }}, {{ outputs.deploy.namespace }}: the deployed release and its namespace.
  • {{ outputs.deploy.revision }}: Helm revision created by the deploy.
  • {{ outputs.deploy.status }}, {{ outputs.verify_release.status }}: Helm status, deployed on success.
  • {{ outputs.deploy.chartVersion }}, {{ outputs.deploy.appVersion }}: chart and application versions now live.
  • {{ outputs.deploy.manifest }}: URI of the manifest Helm actually applied.
  • {{ outputs.rollback.revision }}, {{ outputs.rollback.status }}: where the release ended up after a rollback.

Quick start

  1. Create the five secrets above. In OSS, set them as base64-encoded SECRET_<NAME> environment variables. K8S_CA_CERT is already base64, so it ends up encoded twice in OSS.
  2. Save the flow and run it manually with environment: staging to install the release.
  3. Run it again with environment: production. The run pauses after the dry run. Resume it from the execution page to deploy.
  4. From CI, POST to /api/v1/{tenant}/executions/webhook/company.team/helm-gated-release-with-rollback/{HELM_DEPLOY_WEBHOOK_KEY} with a JSON body, for example curl -X POST -H 'Content-Type: application/json' -d '{"image_tag": "6.14.1"}' https://your-kestra/api/v1/main/executions/webhook/company.team/helm-gated-release-with-rollback/$KEY.
  5. Run helm history podinfo -n podinfo-staging to see the revisions Kestra created, including any rollback.

How to extend

  • Swap podinfo for your own chart: change chart.repository, chart.name and chart.version on the three tasks that use it, or point chart.repository at an oci:// registry.
  • Keep values in Git: wrap the Helm tasks in a io.kestra.plugin.core.flow.WorkingDirectory with io.kestra.plugin.git.Clone and pass the files through valuesFrom.
  • Set cluster and region on the release tasks so Asset metadata matches your naming, e.g. cluster: prod-eu.
  • Use separate connection secrets per environment (for example K8S_MASTER_URL_PRODUCTION) to deploy staging and production to different clusters.
  • Add a smoke test (io.kestra.plugin.core.http.Request against the service) inside release. If it fails, the rollback runs.
  • Prefer rollbackOnFailure: true on deploy if you want Helm to revert by itself. Remove the rollback task in that case, because both together would move the revision twice.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.