Databricks CreateJob

Databricks CreateJob

Certified

Create and run a Databricks job

Creates a Databricks job with one or more tasks, submits it immediately, and optionally waits for completion. Reuse a compute cluster per task with existingClusterId; set waitForCompletion (ISO-8601 duration) to block until the run ends.

yaml
type: io.kestra.plugin.databricks.job.CreateJob

Create a Databricks job, run it, and wait for completion for five minutes.

yaml
id: databricks_job_create
namespace: company.team

tasks:
  - id: create_job
    type: io.kestra.plugin.databricks.job.CreateJob
    authentication:
      token: "{{ secret('DATABRICKS_TOKEN') }}"
    host: "{{ secret('DATABRICKS_HOST') }}"
    jobName: my-databricks-job
    jobTasks:
      - existingClusterId: <your-cluster>
        taskKey: taskKey
        sparkPythonTask:
          pythonFile: /Shared/hello.py
          sparkPythonTaskSource: WORKSPACE
    waitForCompletion: PT5M
Properties

Job name

Name of the Databricks job to create; shown in the Databricks Jobs UI

Min items1

Job tasks

Task definitions; when multiple tasks are present, specify dependsOn for ordering

Definitions
dbtTask
catalogstring

Catalog

Unity Catalog to run the dbt commands against.

commandsarray
SubTypestring

DBT commands

List of dbt commands to execute in order, for example dbt deps then dbt run.

schemastring

Schema

Schema to run the dbt commands against.

warehouseIdstring

SQL warehouse ID

ID of the SQL warehouse used to run the dbt commands.

dependsOnarray
SubTypestring

Task dependencies

List of upstream taskKeys when multiple tasks run in the job

descriptionstring

Task description

existingClusterIdstring

Existing cluster ID

ID of an existing Databricks cluster to run this task on.

librariesarray

Task libraries

Set exactly one of the library types (cran, egg, jar, maven, pypi, or whl).

cran

CRAN library

An R package to install from a CRAN repository.

_packagestring

Package name

Name of the CRAN package to install.

repostring

Repository

CRAN repository URL to install the package from; defaults to the Databricks default repository.

eggstring

Egg library

URI of a Python egg to install (for example a DBFS or cloud storage path).

jarstring

JAR library

URI of a JAR to install (for example a DBFS or cloud storage path).

maven

Maven library

A Maven artifact to install on the cluster.

coordinatesstring

Coordinates

Gradle-style Maven coordinates, for example org.jsoup: jsoup: 1.7.2.

exclusionsarray
SubTypestring

Exclusions

List of dependencies to exclude, for example slf4j: slf4j.

repostring

Repository

Maven repository URL to install the artifact from; defaults to Maven Central.

pypi

PyPI library

A Python package to install from a PyPI repository.

_packagestring

Package name

Name of the PyPI package to install, optionally pinned (for example simplejson==3.8.0).

repostring

Repository

PyPI repository URL to install the package from; defaults to the public PyPI index.

whlstring

Wheel library

URI of a Python wheel (.whl) to install (for example a DBFS or cloud storage path).

notebookTask

Notebook task settings

baseParametersstringobject
SubTypestring

Map of task base parameters.

Can be a map of string/string or a variable that binds to a JSON object.

notebookPathstring

Notebook path

Absolute path of the notebook to run in the Databricks workspace or Git repository.

sourcestring
Possible Values
GITWORKSPACE

Notebook source

Where the notebook lives: WORKSPACE (default) or GIT.

pipelineTask

Pipeline task settings

fullRefreshbooleanstring

Full refresh

If true, the pipeline runs a full refresh, reprocessing all data.

pipelineIdstring

Pipeline ID

ID of the Delta Live Tables pipeline to trigger.

pythonWheelTask

Python Wheel task settings

entryPointstring

Entry point

Named entry point (function or package.module: function) to run from the installed Python wheel.

namedParametersstringobject
SubTypestring

Map of task named parameters.

Can be a map of string/string or a variable that binds to a JSON object.

packageNamestring

Package name

Name of the installed Python wheel package that contains the entry point.

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

runJobTask

Run job task settings

jobIdstring

Job ID

ID of an existing Databricks job to run. Required.

jobParametersstringobject

Job parameters

Map of parameters passed to the triggered job. Can be a map of string/string or a variable that binds to a JSON object.

sparkJarTask

Spark JAR task settings

jarUristring

JAR URI

URI of the JAR to run; the JAR must already be available to the cluster (for example uploaded via a library).

mainClassNamestring

Main class name

Fully qualified name of the class containing the main method to execute.

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

sparkPythonTask

Spark Python task settings

pythonFile*string

Python file

Path or URI of the Python file to run (for example a workspace path or a cloud storage/DBFS URI). Required.

sparkPythonTaskSource*string
Possible Values
GITWORKSPACE

Python file source

Where the Python file lives: WORKSPACE or GIT. Required.

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

sparkSubmitTask

Spark Submit task settings

parametersstringarray

List of task parameters.

Can be a list of strings or a variable that binds to a JSON array of strings.

sqlTask

SQL task settings

parametersstringobject
SubTypestring

Map of task parameters.

Can be a map of string/string or a variable that binds to a JSON object.

queryIdstring

Query ID

ID of an existing saved query in Databricks SQL to run.

warehouseIdstring

SQL warehouse ID

ID of the Databricks SQL warehouse used to run the query.

taskKeystring

Task key

Unique key per task; required when multiple tasks are defined

timeoutSecondsintegerstring

Task timeout (seconds)

Databricks account identifier

Assets this task consumes as inputs or produces as outputs, for lineage tracking and the asset graph (Enterprise Edition). A flow declaring this property on a task is rejected in the open-source edition.

Definitions
assetFailureBehaviorstring
Possible Values
IGNOREFAILWARN

Asset failure behavior

Behavior applied to the task state when a declared asset fails to render, emit, or be persisted (e.g. a lock conflict): FAIL escalates it to FAILED, WARN (default) warns it if it would otherwise succeed, IGNORE leaves the state untouched.

enableAutobooleanstring

Whether to auto-register assets referenced dynamically at runtime that are not statically declared in inputs or outputs.

inputsarray

The assets consumed as inputs.

id*string
Min length1
typestring
outputs

The assets produced as outputs.

id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*object
descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150
id*string
Min length1
Max length150
type*string
Min length1

Custom asset type

descriptionstring
displayNamestring
metadataobject
Default{}
namespacestring
Min length1
Max length150

Databricks authentication configuration

This property allows to configure the authentication to Databricks, different properties should be set depending on the type of authentication and the cloud provider. All configuration options can also be set using the standard Databricks environment variables. Check the Databricks authentication guide for more information.

Definitions
authTypestring

Authentication type

azureClientIdstring

Azure client ID

azureClientSecretstring

Azure client secret

azureTenantIdstring

Azure tenant ID

clientIdstring

Client ID

clientSecretstring

Client secret

googleCredentialsstring

Google credentials JSON

googleServiceAccountstring

Google service account email

passwordstring

Password

tokenstring

Databricks personal access token

usernamestring

Username

Databricks configuration file, use this if you don't want to configure each Databricks account properties one by one

Databricks host

Wait for completion

If set, waits up to the given duration (e.g., PT1H) for the submitted run to finish

Formatduration

Duration

The total run duration; only set when the run has terminated

Formatdate-time

End time

When the run finished executing; only set when the run has terminated

Job identifier

Formaturi

Job console URI

Life cycle state

Set once the run has been submitted; only reaches a terminal value (e.g. TERMINATED, SKIPPED) when waitForCompletion is used

Result state

The run's terminal result state (e.g. SUCCESS, FAILED, TIMEDOUT); only set when the run has terminated

Run identifier

Formaturi

Run console URI

Formatdate-time

Start time

When the run started executing; only set when the run has started

State message

A human-readable description of the run's current state, useful to diagnose a non-SUCCESS result state

The duration of the Databricks run, only available when waitForCompletion is set