
Databricks SubmitRun
CertifiedSubmit a Databricks run
Databricks SubmitRun
Submit a Databricks run
Submits one or more tasks as an ad-hoc run; optionally waits up to waitForCompletion for terminal state. The submission is idempotent: if the Kestra worker running this task is lost and the task is resubmitted, the already in-flight or already completed Databricks run is adopted instead of a duplicate run being launched. A plain task retry after a failure still creates a genuinely new run.
type: io.kestra.plugin.databricks.job.SubmitRunExamples
Submit a Databricks run and wait up to 5 minutes for its completion. A worker-loss resubmit adopts the same Databricks run instead of launching a duplicate one.
id: databricks_job_submit_run
namespace: company.team
tasks:
- id: submit_run
type: io.kestra.plugin.databricks.job.SubmitRun
host: "{{ secret('DATABRICKS_HOST') }}"
authentication:
token: "{{ secret('DATABRICKS_TOKEN') }}"
runTasks:
- existingClusterId: <your-cluster>
taskKey: pysparkTask
sparkPythonTask:
pythonFile: /Shared/hello.py
sparkPythonTaskSource: WORKSPACE
waitForCompletion: PT5M
Properties
runTasks *array
1Run tasks
Task definitions for this run; set dependsOn when multiple tasks are present
io.kestra.plugin.databricks.job.SubmitRun-RunSubmitTaskSetting
Task dependencies
List of upstream taskKeys when multiple tasks run in the same submission
Existing cluster ID
ID of an existing Databricks cluster to run this task on.
Task libraries
Library to install on the cluster
Set exactly one of the library types (cran, egg, jar, maven, pypi, or whl).
CRAN library
An R package to install from a CRAN repository.
io.kestra.plugin.databricks.job.task.LibrarySetting-CranSetting
Package name
Name of the CRAN package to install.
Repository
CRAN repository URL to install the package from; defaults to the Databricks default repository.
Egg library
URI of a Python egg to install (for example a DBFS or cloud storage path).
JAR library
URI of a JAR to install (for example a DBFS or cloud storage path).
Maven library
A Maven artifact to install on the cluster.
io.kestra.plugin.databricks.job.task.LibrarySetting-MavenSetting
Coordinates
Gradle-style Maven coordinates, for example org.jsoup: jsoup: 1.7.2.
Exclusions
List of dependencies to exclude, for example slf4j: slf4j.
Repository
Maven repository URL to install the artifact from; defaults to Maven Central.
PyPI library
A Python package to install from a PyPI repository.
io.kestra.plugin.databricks.job.task.LibrarySetting-PypiSetting
Package name
Name of the PyPI package to install, optionally pinned (for example simplejson==3.8.0).
Repository
PyPI repository URL to install the package from; defaults to the public PyPI index.
Wheel library
URI of a Python wheel (.whl) to install (for example a DBFS or cloud storage path).
Notebook task settings
io.kestra.plugin.databricks.job.task.NotebookTaskSetting
Map of task base parameters.
Can be a map of string/string or a variable that binds to a JSON object.
Notebook path
Absolute path of the notebook to run in the Databricks workspace or Git repository.
GITWORKSPACENotebook source
Where the notebook lives: WORKSPACE (default) or GIT.
Pipeline task settings
Delta Live Tables pipeline task settings
Full refresh
If true, the pipeline runs a full refresh, reprocessing all data.
Pipeline ID
ID of the Delta Live Tables pipeline to trigger.
Python Wheel task settings
io.kestra.plugin.databricks.job.task.PythonWheelTaskSetting
Entry point
Named entry point (function or package.module: function) to run from the installed Python wheel.
Map of task named parameters.
Can be a map of string/string or a variable that binds to a JSON object.
Package name
Name of the installed Python wheel package that contains the entry point.
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
Run job task settings
Run-job task settings
Job ID
ID of an existing Databricks job to run. Required.
Job parameters
Map of parameters passed to the triggered job. Can be a map of string/string or a variable that binds to a JSON object.
Spark JAR task settings
io.kestra.plugin.databricks.job.task.SparkJarTaskSetting
JAR URI
URI of the JAR to run; the JAR must already be available to the cluster (for example uploaded via a library).
Main class name
Fully qualified name of the class containing the main method to execute.
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
Spark Python task settings
io.kestra.plugin.databricks.job.task.SparkPythonTaskSetting
Python file
Path or URI of the Python file to run (for example a workspace path or a cloud storage/DBFS URI). Required.
GITWORKSPACEPython file source
Where the Python file lives: WORKSPACE or GIT. Required.
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
Spark Submit task settings
io.kestra.plugin.databricks.job.task.SparkSubmitTaskSetting
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
Task key
Unique key for this task; required when multiple tasks are defined so dependsOn can reference it.
Task timeout (seconds)
accountId string
Databricks account identifier
assets
Assets this task consumes as inputs or produces as outputs, for lineage tracking and the asset graph (Enterprise Edition). A flow declaring this property on a task is rejected in the open-source edition.
io.kestra.core.models.assets.AssetsDeclaration
IGNOREFAILWARNAsset failure behavior
Behavior applied to the task state when a declared asset fails to render, emit, or be persisted (e.g. a lock conflict): FAIL escalates it to FAILED, WARN (default) warns it if it would otherwise succeed, IGNORE leaves the state untouched.
Whether to auto-register assets referenced dynamically at runtime that are not statically declared in inputs or outputs.
The assets consumed as inputs.
io.kestra.core.models.assets.AssetIdentifier
1The assets produced as outputs.
io.kestra.plugin.ee.assets.Dataset
1150{}1150io.kestra.plugin.ee.assets.File
1150{}1150io.kestra.plugin.ee.assets.Table
1150{}1150io.kestra.plugin.ee.assets.VM
1150{}1150io.kestra.core.models.assets.External
1150{}1150io.kestra.core.models.assets.Custom
11501Custom asset type
{}1150authentication
Databricks authentication configuration
This property allows to configure the authentication to Databricks, different properties should be set depending on the type of authentication and the cloud provider. All configuration options can also be set using the standard Databricks environment variables. Check the Databricks authentication guide for more information.
io.kestra.plugin.databricks.AbstractTask-AuthenticationConfig
Authentication type
Azure client ID
Azure client secret
Azure tenant ID
Client ID
Client secret
Google credentials JSON
Google service account email
Password
Databricks personal access token
Username
configFile string
Databricks configuration file, use this if you don't want to configure each Databricks account properties one by one
host string
Databricks host
idempotencyToken string
Idempotency token seed
Seed used to derive the Databricks idempotency token attached to the run submission, so that a worker-loss resubmit adopts the already in-flight or already completed run instead of launching a duplicate one. Defaults to this task run's Kestra identifier, which is unique per task execution attempt: a plain Kestra retry after a failed run still creates a new Databricks run, while a resubmit of the same attempt after a worker crash adopts the original run. Set this only if you need to key deduplication on something other than the task run itself. Two different executions sharing the same override value will cause the second one to adopt the first one's run.
runName string
Run name
waitForCompletion string
Wait for completion
If set, waits up to the given duration (e.g., PT30M) for the run to finish
Outputs
duration string
durationDuration
The total run duration; only set when the run has terminated
endTime string
date-timeEnd time
When the run finished executing; only set when the run has terminated
lifeCycleState string
Life cycle state
Set once the run has been submitted; only reaches a terminal value (e.g. TERMINATED, SKIPPED) when waitForCompletion is used
resultState string
Result state
The run's terminal result state (e.g. SUCCESS, FAILED, TIMEDOUT); only set when the run has terminated
runId integer
Run identifier
runURI string
uriRun console URI
startTime string
date-timeStart time
When the run started executing; only set when the run has started
stateMessage string
State message
A human-readable description of the run's current state, useful to diagnose a non-SUCCESS result state
Metrics
run.duration timer
The duration of the Databricks run, only available when waitForCompletion is set