Schedule icon
Commands icon
Docker icon
OutputValues icon
Log icon
If icon
Set icon
Fail icon
PurgeCurrentExecutionFiles icon

Renew TLS certificates through ACME, verify them, deploy them and roll back on failure

Renew TLS certificates with ACME and lego in Kestra. Live probe, pre-deploy checks, atomic deploy, strict health check and rollback.

Categories
CloudInfrastructure

Expiry monitors tell you a certificate is about to expire. This flow renews it, and treats the renewal as a deployment: verified before it ships, checked after it ships, and rolled back when clients would break.

It judges every certificate by what the endpoint actually serves, not by a file on disk. A renewed certificate that never got reloaded, a name added to the inventory but not to the certificate, or a chain served without its intermediate all count as broken, because that is what users see.

The demo is complete and local: Pebble, the ACME test server from Let's Encrypt, issues the certificates, its DNS test server answers the DNS-01 challenges, and Caddy is the TLS edge, configured through its admin API. Switching to Let's Encrypt is a change of inputs.

This blueprint was created by zkasuran.

How a certificate is judged

Each domain is probed with its own SNI name. The worst domain decides for the whole certificate.

Status Meaning Renewed
MISSING No certificate served for one of the names, for example a name just added to the inventory yes
EXPIRED Past its end date yes
UNTRUSTED The served chain does not verify against the trusted roots: missing intermediate, wrong CA, staging CA in production yes
SAN_MISMATCH A certificate is served, but it does not cover the name yes
SPLIT The names of one inventory entry are served by different certificates yes
EXPIRING Fewer than renew_before_days left yes
OK Healthy only with force_renew

The checks before and after the deploy

Before anything is deployed, verify_issued requires each new certificate to:

  • chain to a trusted root through the intermediates the CA returned,
  • cover every name in the inventory,
  • match its private key,
  • be valid already.

It warns when a new certificate would already be inside the renewal window, which means the CA issued a shorter certificate than expected. One failure stops the run, and nothing is deployed.

After the deploy, health_check connects again with each SNI name, as a strict client that trusts only the roots and checks the hostname. It requires the new serial on every renewed name, and the unchanged serial on every name that was not renewed, so a deploy cannot break a neighbour. It retries three times, three seconds apart, before it reports failure.

How it works

  1. probe (io.kestra.plugin.scripts.shell.Commands, alpine:3.20 with openssl, curl and jq) runs openssl s_client for every domain and classifies each certificate as above.
  2. plan (io.kestra.plugin.core.output.OutputValues) lists the certificates to renew and keeps what each one was before the run. log_plan prints the plan. With dry_run, the flow stops here.
  3. lifecycle (io.kestra.plugin.core.flow.If) runs the rest only when something is due:
    • renew (goacme/lego:v4.25.2) places one ACME order per certificate with a DNS-01 challenge, then packs the full chains and the keys into issued.tar.
    • verify_issued runs the checks before the deploy.
    • deploy saves the running Caddy configuration as snapshot.json, replaces only the certificates being renewed (matched by tag), keeps every other certificate, and loads the result with one POST /load. Caddy applies it atomically: a rejected configuration leaves the old one running.
    • health_check runs the checks after the deploy. It reports instead of failing, so the next step can act.
    • outcome (io.kestra.plugin.core.flow.If):
      • when healthy, record (io.kestra.plugin.core.kv.Set) merges the new serials, expiry dates and the serials they replaced into the KV entry certificate_register.
      • when not, rollback loads snapshot.json, rollback_check proves every domain serves exactly what it served before the run, and deploy_rolled_back fails the run with the list of problems.
  4. report logs every certificate with its status, and the new serial when it was renewed.
  5. purge_key_material (io.kestra.plugin.core.storage.PurgeCurrentExecutionFiles, in finally) deletes the keys and the snapshot from Kestra's internal storage, whatever the outcome.
  6. daily (io.kestra.plugin.core.trigger.Schedule), disabled by default.

Tested end to end

On Kestra 2.0.3 OSS, with Pebble, pebble-challtestsrv and Caddy 2 in one Docker network. The inventory starts with shop (2 names) and adds api later.

Run Inputs Result
1. Bootstrap only shop, acme_profile: shortlived MISSING, renewed with a 6-day certificate. The verification warns that it is already due. Health check 2/2, register written
2. Dry run full inventory, dry_run shop EXPIRING, 5 day(s) left, api MISSING. Nothing changed
3. Renewal full inventory shop renewed (89 days), api was already issued by an earlier run and stays OK. Health check 3/3
4. Nothing due same Both OK. No order placed
5. Name added api gains api-v2.example.test api MISSING, no certificate served for api-v2.example.test. Only api is reissued with both names. Health check 4/4, shop untouched
6. Broken chain force_renew, demo_drop_intermediate Deployed leaf-only. Each of the 4 names fails as a strict client with unable to get local issuer certificate. Rolled back, rollback check confirms the previous serials on every name, run failed. Register unchanged
7. Wrong trust trust_bundle_url: system Every certificate UNTRUSTED. New certificates fail verify_issued with the same error. Nothing deployed, previous serials still served
8. CA refuses name blocked-domain.example Pebble returns rejectedIdentifier, renew fails, nothing deployed

Run 6 is the reason the health check uses a strict client. Caddy loaded the leaf-only configuration without complaint, and a browser that has seen the intermediate before still connects. Clients that have not, such as other services, mobile apps and curl, fail.

Problems found while testing

  • No profile, random lifetime. Without a profile, Pebble picks one at random. Two identical renewals returned a 90-day and a 6-day certificate. The flow now always requests a profile (default), and verify_issued warns when a new certificate would already be due.
  • An empty input becomes the default. Leaving trust_bundle_url empty did not select the system trust store: Kestra applied the default instead. The input now takes the explicit value system.
  • Sequential DNS challenges wait a minute. The exec DNS provider solves names one after another, waiting the propagation timeout in between. EXEC_SEQUENCE_INTERVAL brings a 2-name order from 70 seconds to about 10.
  • Kestra 2.0 has no flow-level pluginDefaults and no ForEach. The image and runner are set on each task, and the register is one merged KV entry instead of one entry per certificate.

Inputs

Input Default Purpose
certificates shop and api Inventory: name and domains. The name is also the Caddy tag
renew_before_days 30 Renewal window
tls_endpoint web:443 Where the certificates are served
caddy_admin_url http://web:2019 The deploy target
acme_directory Pebble The ACME server
acme_email ops@example.test Account contact
acme_profile default Certificate profile
key_type ec256 ec256, ec384 or RSA
trust_bundle_url Pebble roots system for a public CA
force_renew, dry_run false
demo_drop_intermediate false Demo of a broken chain

Run the demo

Start the three services in the network your Kestra worker uses for task containers, then run the flow. Set networkMode on the Docker task runners if that network is not the default one.

docker network create acme-demo
docker run -d --name challtestsrv --network acme-demo ghcr.io/letsencrypt/pebble-challtestsrv:latest \
  -dnsserver :8053 -management :8055 -http01 "" -https01 "" -tlsalpn01 "" -doh ""
docker run -d --name pebble --network acme-demo -e PEBBLE_VA_NOSLEEP=1 -e PEBBLE_WFE_NONCEREJECT=0 \
  ghcr.io/letsencrypt/pebble:latest -config /test/config/pebble-config.json -dnsserver challtestsrv:8053
cat > caddy.json <<'JSON'
{"admin": {"listen": "0.0.0.0:2019", "origins": ["web:2019"]},
 "apps": {"http": {"servers": {"edge": {"listen": [":443"], "automatic_https": {"disable": true},
   "routes": [{"handle": [{"handler": "static_response", "body": "ok {http.request.host}\n"}]}]}}}}}
JSON
docker run -d --name web --network acme-demo -v "$PWD/caddy.json:/etc/caddy/boot.json:ro" \
  caddy:2-alpine caddy run --config /etc/caddy/boot.json

Then:

  1. Run with only shop and acme_profile: shortlived.
  2. Run with the defaults: shop is now expiring and api is missing.
  3. Run again: nothing to do.
  4. Run with force_renew and demo_drop_intermediate to watch the rollback.

The Caddy admin API has no authentication. Keep it on a private network, as here.

Use it with Let's Encrypt

  • acme_directory: https://acme-staging-v02.api.letsencrypt.org/directory first, then https://acme-v02.api.letsencrypt.org/directory.
  • acme_profile: classic, tlsserver or shortlived.
  • trust_bundle_url: system.
  • In renew:
    • set DNS_PROVIDER to your provider, for example cloudflare, route53, gcloud or azuredns, and add its credentials to env from secrets, for example CLOUDFLARE_DNS_API_TOKEN: "{{ secret('CLOUDFLARE_DNS_API_TOKEN') }}";
    • remove EXEC_*, DNS_RESOLVERS (or point it at a public resolver), LEGO_CA_CERTIFICATES and the two demo files;
    • keep the account key between runs instead of registering each time: mount a volume at ./lego, or restore and save lego/accounts from your secret store.
  • Let's Encrypt rate limits apply: 50 certificates per registered domain per week, 5 duplicate certificates per week. Test with staging.

Use another deploy target

The flow only expects deploy to save a rollback point and apply the change, and rollback to restore it. Some ways to do that:

  • nginx or HAProxy: copy the files over SSH (io.kestra.plugin.fs.ssh.Command), run nginx -t, reload, and keep the previous files for the rollback.
  • Kubernetes: update the TLS secret, after saving the previous one.
  • AWS: aws acm import-certificate --certificate-arn on the existing ARN. Keep the previous certificate for the rollback.

The probe and the health check do not change, because they only look at what the endpoint serves.

Expected outputs

  • outputs.probe.vars.probe: status, reason, serial and days left for each certificate.
  • outputs.verify_issued.vars.certs: serial, expiry, key, issuer and SANs of each new certificate.
  • outputs.health_check.vars: healthy, checks, failures.
  • KV certificate_register: for each certificate, the serial, not_after, SANs, the serial it replaced and why, and when and by which execution it was renewed.

Things to know

  • Private keys pass through Kestra internal storage between renew, verify_issued and deploy, and the Caddy snapshot contains every key it serves. purge_key_material deletes them at the end of each run. They also appear in the Caddy configuration, so protect its admin API.
  • concurrency: limit 1 prevents two runs from deploying at the same time.
  • An endpoint that serves a default certificate for unknown names reports SAN_MISMATCH for a new name. One that serves nothing, like Caddy here, reports MISSING.
  • The demo trusts Pebble's root by fetching it with curl -k. That is only acceptable for a test CA.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.