Configure Retry Strategies for Transient Task Failures

For the complete documentation index, see llms.txt. For a full content snapshot, see llms-full.txt. Append .md to any kestra.io/docs/* URL for plain Markdown.

Retries automatically rerun failed tasks. Each retry creates a new task run attempt based on the retry configuration defined in the flow.

Task-level retries

Example

This task retries up to 5 times with a 15-minute interval between attempts:

- id: retry_sample
type: io.kestra.plugin.core.log.Log
message: my output for task {{ task.id }}
timeout: PT10M
retry:
type: constant
maxAttempts: 5
interval: PT15M

In this example, the flow retries 4 times every 0.25 seconds. It succeeds on the 5th attempt, using {{ taskrun.attemptsCount }} to track retries:

id: retry
namespace: company.team
description: This flow retries 4 times and succeeds on the 5th attempt
tasks:
- id: failed
type: io.kestra.plugin.scripts.shell.Commands
taskRunner:
type: io.kestra.plugin.core.runner.Process
commands:
- 'if [ "{{ taskrun.attemptsCount }}" -eq 4 ]; then exit 0; else exit 1; fi'
retry:
type: constant
interval: PT0.25S
maxAttempts: 5
maxDuration: PT1M
warningOnRetry: true
errors:
- id: never_happen
type: io.kestra.plugin.core.debug.Return
format: "Never happened {{ task.id }}"

Timeout vs. Max Retry Duration

  • timeout: Maximum duration for a single task attempt (initial or retry). If exceeded, the attempt fails.
  • retry.maxDuration: Maximum total time allowed for the task, including all attempts and delays. Once exceeded, retries stop.

Example: With timeout: 10m and maxDuration: 30m:

  • Each attempt can last up to 10 minutes.
  • The overall retries stop after 30 minutes in total.

Retry options

NameTypeDescription
typestringRetry strategy: constant, exponential, or random.
maxAttemptsintegerMaximum number of attempts, including the initial run.
maxDurationDurationMaximum total time for the task, across all attempts.
warningOnRetryBooleanMarks execution as WARNING if retries occurred (default: false).

Duration format

Durations use ISO 8601 format (weeks, months, years not supported). Examples:

ValueDescription
PT0.25S250 ms
PT2S2 seconds
PT1M1 minute
PT3.5H3 hours, 30 minutes
P6DT4H6 days, 4 hours

Retry on flowable tasks

Flowable tasks such as Sequential and Parallel accept a retry block, but it does not behave like a task-level retry. A retry on a flowable task does not rerun the group. Instead, it sets the default retry policy that child tasks inherit when they have no retry of their own.

id: token_and_api
namespace: company.team
tasks:
- id: group
type: io.kestra.plugin.core.flow.Sequential
retry:
type: constant
maxAttempts: 2
interval: PT1S
tasks:
- id: get_token
type: io.kestra.plugin.core.debug.Return
format: "{{ now() }}"
- id: call_api
type: io.kestra.plugin.core.execution.Fail
errorMessage: "Token expired"

If call_api fails, Kestra retries only call_api. get_token does not run again, and each retry of call_api uses the same token from the first run. A token that expires between the two tasks is never refreshed this way.

Retry a group of tasks as one unit

To retry a set of tasks together so that every task in the group reruns on failure, move them into a subflow and place retry on the Subflow task. Each retry creates a new child execution, so all tasks in the subflow run again from the start.

id: parent
namespace: company.team
tasks:
- id: group
type: io.kestra.plugin.core.flow.Subflow
namespace: company.team
flowId: get_token_and_call_api
retry:
type: constant
maxAttempts: 2
interval: PT1S

Token refresh pattern

For the specific case of a short-lived token that may expire between tasks, Credentials (Enterprise Edition and Cloud) is the cleaner solution. Kestra fetches and refreshes the token automatically, so every retry gets a valid one without needing a get_token task at all.

id: api_call
namespace: company.team
tasks:
- id: call_api
type: io.kestra.plugin.core.http.Request
uri: https://api.example.com/v1/data
options:
auth:
type: BEARER
token: "{{ credential('my_oauth') }}"
retry:
type: constant
maxAttempts: 3
interval: PT1S

Retry types

constant

Retries at fixed intervals. Example: with interval: PT10M, retries occur every 10 minutes.

NameTypeDescription
intervalDurationDelay between attempts.

exponential

Wait time increases after each retry (e.g., 30s, 1m, 2m, …).

NameTypeDescription
intervalDurationBase interval between attempts.
maxIntervalDurationMaximum interval allowed.
delayFactorDoubleMultiplier (default: 2). Example: interval 30s → 30s, 1m, 2m, 4m…

random

Randomized delays within bounds.

NameTypeDescription
minIntervalDurationMinimum delay.
maxIntervalDurationMaximum delay.

Configuring retries globally

You can configure retries globally for all tasks in Kestra:

kestra:
plugins:
configurations:
- type: io.kestra
values:
retry:
type: constant
maxAttempts: 3
interval: PT30S

This applies a constant retry policy with up to 3 attempts every 30 seconds.

Flow-level retries

You can retry at the flow level, restarting either the entire execution or just failed tasks. Options:

  1. CREATE_NEW_EXECUTION: Start a new execution.
  2. RETRY_FAILED_TASK: Retry only the failed task.
id: flow_level_retry
namespace: company.team
retry:
maxAttempts: 3
behavior: CREATE_NEW_EXECUTION # or RETRY_FAILED_TASK
type: constant
interval: PT1S
tasks:
- id: fail_1
type: io.kestra.plugin.core.execution.Fail
allowFailure: true
- id: fail_2
type: io.kestra.plugin.core.execution.Fail
  • With CREATE_NEW_EXECUTION, the execution attempt increases.
  • With RETRY_FAILED_TASK, only the task run attempt increases.

Retry vs. Restart vs. Replay

Retry is the only automatic mechanism. Restart and Replay are both manual, initiated from the UI.

ConceptScopeTriggerNew execution?
RetryTask levelAutomaticNo
RestartFlow levelManualNo
ReplayFlow or task levelManualYes

Restart reruns only the failed tasks within the same execution, keeping the same execution ID. Use the Restart button at the top of the Execution overview page.

Replay starts a new execution from any task — successful or failed — and assigns it a new execution ID. Previous task outputs are reused from cache when available. Trigger a replay from the Actions menu at the top of the Execution overview page, or directly from a task node in the Topology, Gantt, or Logs view. See the Replay documentation.

After a replay, the new execution’s Overview tab shows an Original Execution field linking back to the source run.

Was this page helpful?