Databricks SubmitRun

Databricks SubmitRun

Certified

Submit a Databricks run

Submits one or more tasks as an ad-hoc run; optionally waits up to waitForCompletion for terminal state. The submission is idempotent: if the Kestra worker running this task is lost and the task is resubmitted, the already in-flight or already completed Databricks run is adopted instead of a duplicate run being launched. A plain task retry after a failure still creates a genuinely new run.

yaml
type: io.kestra.plugin.databricks.job.SubmitRun

Submit a Databricks run and wait up to 5 minutes for its completion. A worker-loss resubmit adopts the same Databricks run instead of launching a duplicate one.

yaml
id: databricks_job_submit_run
namespace: company.team

tasks:
  - id: submit_run
    type: io.kestra.plugin.databricks.job.SubmitRun
    host: "{{ secret('DATABRICKS_HOST') }}"
    authentication:
      token: "{{ secret('DATABRICKS_TOKEN') }}"
    runTasks:
      - existingClusterId: <your-cluster>
        taskKey: pysparkTask
        sparkPythonTask:
          pythonFile: /Shared/hello.py
          sparkPythonTaskSource: WORKSPACE
    waitForCompletion: PT5M
Properties
Min items1

Run tasks

Task definitions for this run; set dependsOn when multiple tasks are present

Definitions
dependsOnarray
SubTypestring

Task dependencies

List of upstream taskKeys when multiple tasks run in the same submission

existingClusterIdstring

Existing cluster ID

ID of an existing Databricks cluster to run this task on.

librariesarray

Task libraries

Set exactly one of the library types (cran, egg, jar, maven, pypi, or whl).

cran

CRAN library

An R package to install from a CRAN repository.

_packagestring

Package name

Name of the CRAN package to install.

repostring

Repository

CRAN repository URL to install the package from; defaults to the Databricks default repository.

eggstring

Egg library

URI of a Python egg to install (for example a DBFS or cloud storage path).

jarstring

JAR library

URI of a JAR to install (for example a DBFS or cloud storage path).

maven

Maven library

A Maven artifact to install on the cluster.

coordinatesstring

Coordinates

Gradle-style Maven coordinates, for example org.jsoup: jsoup: 1.7.2.

exclusionsarray
SubTypestring

Exclusions

List of dependencies to exclude, for example slf4j: slf4j.

repostring

Repository

Maven repository URL to install the artifact from; defaults to Maven Central.

pypi

PyPI library

A Python package to install from a PyPI repository.

_packagestring

Package name

Name of the PyPI package to install, optionally pinned (for example simplejson==3.8.0).

repostring

Repository

PyPI repository URL to install the package from; defaults to the public PyPI index.

whlstring

Wheel library

URI of a Python wheel (.whl) to install (for example a DBFS or cloud storage path).

notebookTask

Notebook task settings

baseParametersstringobject
SubTypestring

Map of task base parameters.

Can be a map of string/string or a variable that binds to a JSON object.

notebookPathstring

Notebook path

Absolute path of the notebook to run in the Databricks workspace or Git repository.

sourcestring
Possible Values
GITWORKSPACE

Notebook source

Where the notebook lives: WORKSPACE (default) or GIT.

pipelineTask

Pipeline task settings

fullRefreshbooleanstring

Full refresh

If true, the pipeline runs a full refresh, reprocessing all data.

pipelineIdstring

Pipeline ID

ID of the Delta Live Tables pipeline to trigger.

pythonWheelTask

Python Wheel task settings

entryPointstring

Entry point

Named entry point (function or package.module: function) to run from the installed Python wheel.

namedParametersstringobject
SubTypestring

Map of task named parameters.

Can be a map of string/string or a variable that binds to a JSON object.

packageNamestring

Package name

Name of the installed Python wheel package that contains the entry point.

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

runJobTask

Run job task settings

jobIdstring

Job ID

ID of an existing Databricks job to run. Required.

jobParametersstringobject

Job parameters

Map of parameters passed to the triggered job. Can be a map of string/string or a variable that binds to a JSON object.

sparkJarTask

Spark JAR task settings

jarUristring

JAR URI

URI of the JAR to run; the JAR must already be available to the cluster (for example uploaded via a library).

mainClassNamestring

Main class name

Fully qualified name of the class containing the main method to execute.

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

sparkPythonTask

Spark Python task settings

pythonFile*string

Python file

Path or URI of the Python file to run (for example a workspace path or a cloud storage/DBFS URI). Required.

sparkPythonTaskSource*string
Possible Values
GITWORKSPACE

Python file source

Where the Python file lives: WORKSPACE or GIT. Required.

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

sparkSubmitTask

Spark Submit task settings

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

taskKeystring

Task key

Unique key for this task; required when multiple tasks are defined so dependsOn can reference it.

timeoutSecondsinteger

Task timeout (seconds)

Databricks account identifier

Assets this task consumes as inputs or produces as outputs, for lineage tracking and the asset graph (Enterprise Edition). A flow declaring this property on a task is rejected in the open-source edition.

Definitions
assetFailureBehaviorstring
Possible Values
IGNOREFAILWARN

Asset failure behavior

Behavior applied to the task state when a declared asset fails to render, emit, or be persisted (e.g. a lock conflict): FAIL escalates it to FAILED, WARN (default) warns it if it would otherwise succeed, IGNORE leaves the state untouched.

enableAutobooleanstring

Whether to auto-register assets referenced dynamically at runtime that are not statically declared in inputs or outputs.

inputsarray

The assets consumed as inputs.

id*string
Min length1
typestring
outputs

The assets produced as outputs.

id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*string
Min length1

Custom asset type

descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150

Databricks authentication configuration

This property allows to configure the authentication to Databricks, different properties should be set depending on the type of authentication and the cloud provider. All configuration options can also be set using the standard Databricks environment variables. Check the Databricks authentication guide for more information.

Definitions
authTypestring

Authentication type

azureClientIdstring

Azure client ID

azureClientSecretstring

Azure client secret

azureTenantIdstring

Azure tenant ID

clientIdstring

Client ID

clientSecretstring

Client secret

googleCredentialsstring

Google credentials JSON

googleServiceAccountstring

Google service account email

passwordstring

Password

tokenstring

Databricks personal access token

usernamestring

Username

Databricks configuration file, use this if you don't want to configure each Databricks account properties one by one

Databricks host

Idempotency token seed

Seed used to derive the Databricks idempotency token attached to the run submission, so that a worker-loss resubmit adopts the already in-flight or already completed run instead of launching a duplicate one. Defaults to this task run's Kestra identifier, which is unique per task execution attempt: a plain Kestra retry after a failed run still creates a new Databricks run, while a resubmit of the same attempt after a worker crash adopts the original run. Set this only if you need to key deduplication on something other than the task run itself. Two different executions sharing the same override value will cause the second one to adopt the first one's run.

Run name

Wait for completion

If set, waits up to the given duration (e.g., PT30M) for the run to finish

Formatduration

Duration

The total run duration; only set when the run has terminated

Formatdate-time

End time

When the run finished executing; only set when the run has terminated

Life cycle state

Set once the run has been submitted; only reaches a terminal value (e.g. TERMINATED, SKIPPED) when waitForCompletion is used

Result state

The run's terminal result state (e.g. SUCCESS, FAILED, TIMEDOUT); only set when the run has terminated

Run identifier

Formaturi

Run console URI

Formatdate-time

Start time

When the run started executing; only set when the run has started

State message

A human-readable description of the run's current state, useful to diagnose a non-SUCCESS result state

The duration of the Databricks run, only available when waitForCompletion is set