Commands icon
Docker icon
Log icon
SlackIncomingWebhook icon
Schedule icon

Migrate a Cron Script to Kestra Without Rewriting It, Using the Strangler Fig Pattern

Wrap an existing Python script in Kestra with a Docker volume mount, secrets, a timeout and a concurrency limit, then run it in parallel with cron before cutting over.

Categories
CoreData

Teams stall on orchestration migrations because they assume migration means rewriting. It does not. This blueprint is the wrapper described in Kai Waehner's write-up of moving a working Python script off cron: a Kestra flow starts a container, mounts the script's directory into it, injects secrets as environment variables, and runs the script untouched. Kestra neither knows nor cares what happens inside. The contract is an exit code and a directory, which is exactly why the script needs no changes and why the team that owns it does not need to learn Kestra.

The worked example is a real data pipeline: the wrapped script collects new posts from more than a hundred tech blogs, ranks them with an LLM, and sends one digest email. That shape matters, because the scraping, the ranking and the state file all stay inside the script where they already work. Kestra takes over scheduling, secrets, isolation and observability without ever parsing a feed itself. If you want the same digest built natively out of Kestra tasks instead, the RSS news digest blueprint does that; this one is for the case where the pipeline already exists and you do not want to rewrite it.

The migration follows the strangler fig pattern in four phases: wrap the script unchanged, run it in parallel with cron, cut over, and only then, if ever, refactor. This flow is phase one, and it ships defaulted to dry-run so the parallel phase is safe by default.

How it works

  1. The run_script task (io.kestra.plugin.scripts.python.Commands) runs the script inside python:3.13-slim through the Docker task runner, with volumes mounting work_dir at /work. The script is executed from the mount, not copied in, so it keeps its own state file on disk.
  2. beforeCommands installs the script's dependencies at run time, which avoids maintaining a custom image on day one.
  3. Secrets arrive as env entries from Kestra's secret store. A script whose loader uses os.environ.setdefault will prefer these over any local .env file, so credential delivery moves to Kestra with no code change.
  4. The mode input switches the script between --dry-run and its real behaviour, using Pebble's ternary. Note that Kestra has no ?: shorthand, so the full condition ? a : b form is required.
  5. timeout: PT15M is the part that earns its keep. Cron reports neither runtime nor failure, so a run that quietly takes ten times longer than normal looks identical to a healthy one. A timeout turns slow into failed loudly.
  6. Flow-level concurrency with limit: 1 and behavior: QUEUE guarantees two executions can never race on the same state file.
  7. The errors block alerts on failure using a separate webhook secret, because an alert channel must not share a failure mode with the thing it watches.

What you get

  • A working script under orchestration on day one, with zero lines of the script changed.
  • A schedule with a real timezone, a timeout, and a concurrency limit, all as configuration rather than code.
  • Encrypted secrets with role-based access instead of file permissions and a .env on disk.
  • Run history with logs, durations, and exit codes, which is the signal cron never gave you.
  • A clean split of ownership: the script team keeps writing Python, the platform team owns the flow.

Who it's for

  • Platform teams inheriting a pile of cron jobs nobody has time to rewrite.
  • Script owners who want observability without adopting a framework.
  • Anyone running a scheduled job whose only failure signal today is a human noticing that an email did not arrive.

Why orchestrate this with Kestra

The argument for wrapping rather than rewriting is that the day-one cost is close to zero and every later improvement has somewhere to land. Cron gives you a schedule and nothing else: no runtime, no history, no retry, no alert, no way to stop two runs colliding on a state file. Kestra adds all of those as configuration around a script it never has to understand. The visibility alone tends to pay for the migration, because a duration on a screen surfaces the class of bug that never fails and therefore never gets noticed.

Prerequisites

  • A Kestra worker with access to the Docker socket, and volume mounting enabled for the Docker task runner.
  • The script and its state on a path the worker can reach, given in work_dir.
  • A script that derives its paths from its own location rather than the working directory. This is what lets you point the flow at a parallel copy during the migration.
  • A Slack incoming webhook for failure alerts, ideally not the one the script itself posts to.

Secrets

  • ANTHROPIC_API_KEY, SCRIPT_PASSWORD: rename these to whatever your script reads. They are passed straight through as environment variables.
  • ALERT_SLACK_WEBHOOK: webhook for the failure alert, deliberately separate from anything the script uses.

Quick start

  1. Point work_dir at a copy of the script directory, not production, and leave mode on dry-run.
  2. Replace the env entries and beforeCommands with your script's real credentials and dependencies.
  3. Set the cron and timezone to match the crontab entry you are replacing, and leave the cron job running.
  4. Compare the two for a few days. You are looking at durations as much as outcomes.
  5. To cut over: comment out the crontab line, repoint work_dir at the production directory, and switch mode to send. Nothing is copied, so there is no gap and no overlap.

How to extend

  • Take a lock on the state file itself with io.kestra.plugin.kestra.ee.locks.Acquire and a Release in a finally block, so exactly one writer is a property of the file rather than of this flow.
  • Open a Case on failure instead of only sending a message, so a recurring problem accumulates history against one deduplicated incident with an owner and an SLA.
  • Add an external heartbeat pinged after every successful run, since an errors block cannot fire for an execution that never started.
  • Decompose the script's internal loop into per-task iterations once it is stable, at which point the execution timeline shows exactly which step is slow.
  • Move flow deployment to Git sync so a production change is a reviewed merge.
  • Swap the Schedule trigger for an event-driven trigger when the work stops being time-based.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.