
Databricks CreateJob
CertifiedCreate and run a Databricks job
Databricks CreateJob
Create and run a Databricks job
Creates a Databricks job with one or more tasks, submits it immediately, and optionally waits for completion. Reuse a compute cluster per task with existingClusterId; set waitForCompletion (ISO-8601 duration) to block until the run ends.
type: io.kestra.plugin.databricks.job.CreateJobExamples
Create a Databricks job, run it, and wait for completion for five minutes.
id: databricks_job_create
namespace: company.team
tasks:
- id: create_job
type: io.kestra.plugin.databricks.job.CreateJob
authentication:
token: "{{ secret('DATABRICKS_TOKEN') }}"
host: "{{ secret('DATABRICKS_HOST') }}"
jobName: my-databricks-job
jobTasks:
- existingClusterId: <your-cluster>
taskKey: taskKey
sparkPythonTask:
pythonFile: /Shared/hello.py
sparkPythonTaskSource: WORKSPACE
waitForCompletion: PT5M
Properties
jobName *string
Job name
Name of the Databricks job to create; shown in the Databricks Jobs UI
jobTasks *array
1Job tasks
Task definitions; when multiple tasks are present, specify dependsOn for ordering
io.kestra.plugin.databricks.job.CreateJob-JobTaskSetting
DBT task settings
Catalog
Unity Catalog to run the dbt commands against.
DBT commands
List of dbt commands to execute in order, for example dbt deps then dbt run.
Schema
Schema to run the dbt commands against.
SQL warehouse ID
ID of the SQL warehouse used to run the dbt commands.
Task dependencies
List of upstream taskKeys when multiple tasks run in the job
Task description
Existing cluster ID
ID of an existing Databricks cluster to run this task on.
Task libraries
Library to install on the cluster
Set exactly one of the library types (cran, egg, jar, maven, pypi, or whl).
CRAN library
An R package to install from a CRAN repository.
io.kestra.plugin.databricks.job.task.LibrarySetting-CranSetting
Package name
Name of the CRAN package to install.
Repository
CRAN repository URL to install the package from; defaults to the Databricks default repository.
Egg library
URI of a Python egg to install (for example a DBFS or cloud storage path).
JAR library
URI of a JAR to install (for example a DBFS or cloud storage path).
Maven library
A Maven artifact to install on the cluster.
io.kestra.plugin.databricks.job.task.LibrarySetting-MavenSetting
Coordinates
Gradle-style Maven coordinates, for example org.jsoup: jsoup: 1.7.2.
Exclusions
List of dependencies to exclude, for example slf4j: slf4j.
Repository
Maven repository URL to install the artifact from; defaults to Maven Central.
PyPI library
A Python package to install from a PyPI repository.
io.kestra.plugin.databricks.job.task.LibrarySetting-PypiSetting
Package name
Name of the PyPI package to install, optionally pinned (for example simplejson==3.8.0).
Repository
PyPI repository URL to install the package from; defaults to the public PyPI index.
Wheel library
URI of a Python wheel (.whl) to install (for example a DBFS or cloud storage path).
Notebook task settings
io.kestra.plugin.databricks.job.task.NotebookTaskSetting
Map of task base parameters.
Can be a map of string/string or a variable that binds to a JSON object.
Notebook path
Absolute path of the notebook to run in the Databricks workspace or Git repository.
GITWORKSPACENotebook source
Where the notebook lives: WORKSPACE (default) or GIT.
Pipeline task settings
Delta Live Tables pipeline task settings
Full refresh
If true, the pipeline runs a full refresh, reprocessing all data.
Pipeline ID
ID of the Delta Live Tables pipeline to trigger.
Python Wheel task settings
io.kestra.plugin.databricks.job.task.PythonWheelTaskSetting
Entry point
Named entry point (function or package.module: function) to run from the installed Python wheel.
Map of task named parameters.
Can be a map of string/string or a variable that binds to a JSON object.
Package name
Name of the installed Python wheel package that contains the entry point.
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
Run job task settings
Run-job task settings
Job ID
ID of an existing Databricks job to run. Required.
Job parameters
Map of parameters passed to the triggered job. Can be a map of string/string or a variable that binds to a JSON object.
Spark JAR task settings
io.kestra.plugin.databricks.job.task.SparkJarTaskSetting
JAR URI
URI of the JAR to run; the JAR must already be available to the cluster (for example uploaded via a library).
Main class name
Fully qualified name of the class containing the main method to execute.
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
Spark Python task settings
io.kestra.plugin.databricks.job.task.SparkPythonTaskSetting
Python file
Path or URI of the Python file to run (for example a workspace path or a cloud storage/DBFS URI). Required.
GITWORKSPACEPython file source
Where the Python file lives: WORKSPACE or GIT. Required.
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
Spark Submit task settings
io.kestra.plugin.databricks.job.task.SparkSubmitTaskSetting
List of task parameters.
Can be a list of strings or a variable that binds to a JSON array of strings.
SQL task settings
io.kestra.plugin.databricks.job.task.SqlTaskSetting
Map of task parameters.
Can be a map of string/string or a variable that binds to a JSON object.
Query ID
ID of an existing saved query in Databricks SQL to run.
SQL warehouse ID
ID of the Databricks SQL warehouse used to run the query.
Task key
Unique key per task; required when multiple tasks are defined
Task timeout (seconds)
accountId string
Databricks account identifier
assets
Assets this task consumes as inputs or produces as outputs, for lineage tracking and the asset graph (Enterprise Edition). A flow declaring this property on a task is rejected in the open-source edition.
io.kestra.core.models.assets.AssetsDeclaration
IGNOREFAILWARNAsset failure behavior
Behavior applied to the task state when a declared asset fails to render, emit, or be persisted (e.g. a lock conflict): FAIL escalates it to FAILED, WARN (default) warns it if it would otherwise succeed, IGNORE leaves the state untouched.
Whether to auto-register assets referenced dynamically at runtime that are not statically declared in inputs or outputs.
The assets consumed as inputs.
io.kestra.core.models.assets.AssetIdentifier
1The assets produced as outputs.
io.kestra.plugin.ee.assets.Dataset
1150{}1150io.kestra.plugin.ee.assets.File
1150{}1150io.kestra.plugin.ee.assets.Table
1150{}1150io.kestra.plugin.ee.assets.VM
1150{}1150io.kestra.core.models.assets.External
1150{}1150io.kestra.core.models.assets.Custom
11501Custom asset type
{}1150authentication
Databricks authentication configuration
This property allows to configure the authentication to Databricks, different properties should be set depending on the type of authentication and the cloud provider. All configuration options can also be set using the standard Databricks environment variables. Check the Databricks authentication guide for more information.
io.kestra.plugin.databricks.AbstractTask-AuthenticationConfig
Authentication type
Azure client ID
Azure client secret
Azure tenant ID
Client ID
Client secret
Google credentials JSON
Google service account email
Password
Databricks personal access token
Username
configFile string
Databricks configuration file, use this if you don't want to configure each Databricks account properties one by one
host string
Databricks host
waitForCompletion string
Wait for completion
If set, waits up to the given duration (e.g., PT1H) for the submitted run to finish
Outputs
duration string
durationDuration
The total run duration; only set when the run has terminated
endTime string
date-timeEnd time
When the run finished executing; only set when the run has terminated
jobId integer
Job identifier
jobURI string
uriJob console URI
lifeCycleState string
Life cycle state
Set once the run has been submitted; only reaches a terminal value (e.g. TERMINATED, SKIPPED) when waitForCompletion is used
resultState string
Result state
The run's terminal result state (e.g. SUCCESS, FAILED, TIMEDOUT); only set when the run has terminated
runId integer
Run identifier
runURI string
uriRun console URI
startTime string
date-timeStart time
When the run started executing; only set when the run has started
stateMessage string
State message
A human-readable description of the run's current state, useful to diagnose a non-SUCCESS result state
Metrics
run.duration timer
The duration of the Databricks run, only available when waitForCompletion is set