Schedule icon
CreateProcessInstance icon
If icon
Log icon
Switch icon
SlackIncomingWebhook icon

Camunda Process Outcome Compliance Gate

Run a Camunda 8 process, classify completion by its outcome variable, alert on risk, and optionally retrigger a fresh run.

Categories
BusinessInfrastructure

Camunda Operate shows whether a process instance is active, completed, or has an incident, but nothing pages anyone when a business process technically finishes without ever reaching the approved path - a validation branch quietly routes every order to a rejected end event and nobody notices until finance asks why nothing shipped. This blueprint uses io.kestra.plugin.camunda.CreateProcessInstance with awaitCompletion: true to run a named process and read back the one variable that is supposed to record its outcome, classifies the result into a three-state risk level, alerts on anything but the expected outcome, and can start one fresh retry instance when the run was not healthy, behind two independent gates.

How it works

  1. periodic_outcome_check (io.kestra.plugin.core.trigger.Schedule) runs every hour, always forcing auto_remediate: "false" and dry_run: "true" through its own inputs: override regardless of this flow's defaults. Shipped disabled so you can validate a manual run first.
  2. run_monitored_process (io.kestra.plugin.camunda.CreateProcessInstance, awaitCompletion: true, fetchVariables: [expected_outcome_variable]) starts process_id and blocks until it ends, returning processInstanceKey and a variables map containing only the one variable asked for.
  3. evaluate_outcome_presence (io.kestra.plugin.core.flow.If) branches on whether variables[expected_outcome_variable] is defined at all.
    • else (UNREACHABLE): alert_outcome_unreachable reports that the expected variable never came back from a completed instance.
    • then (variable present): nested evaluate_outcome_value (io.kestra.plugin.core.flow.If) branches on whether that variable equals expected_outcome_value.
      • then (HEALTHY): log_outcome_healthy records a clean, expected outcome.
      • else (AT_RISK): route_remediation (io.kestra.plugin.core.flow.Switch on auto_remediate) either considers a retry ("true", gated again by dry_run through check_dry_run) or only logs that remediation is disabled ("false"). When both gates allow it, retry_process_instance starts and awaits one fresh instance of the same process_id. alert_outcome_at_risk always fires to Slack regardless of which path ran.
  4. log_audit_complete always runs last, printing the headline numbers regardless of which branch fired.
  5. The errors block alerts Slack separately if the flow itself fails outright - an unreachable cluster, bad credentials, or the instance never reaching an end event within requestTimeout must not read as "no risk".

What you get

  • A direct read of the business outcome a BPMN process itself recorded, not an inference from process-instance state alone (Camunda's own client for this task returns no separate success/failure flag - only whatever variables the model sets).
  • A distinct UNREACHABLE state so a process that ends without ever setting the expected variable (a modeling gap) is never conflated with one that explicitly recorded an unwanted outcome.
  • Remediation that only ever starts a brand new instance - it never cancels, modifies, or re-reads the original one.
  • Two independent gates (auto_remediate, dry_run) between "unexpected outcome" and "a retry actually started".
  • A final status log every run, risk or not, so the execution history doubles as an outcome trend for that process.

Who it's for

  • Process owners and platform teams running Camunda 8 who need an automated check that a critical process is actually reaching its intended end state, not just completing without error.
  • Teams who model business outcomes as a process variable (approval status, validation result, routing decision) and want that variable treated as a first-class compliance signal.
  • Anyone evaluating the Camunda plugin who wants a worked example of CreateProcessInstance with awaitCompletion driving both a monitored run and a separate retry call, since this is the only plugin in this repo's governance-gate series built on a BPM/workflow engine rather than a database or message broker.

Why orchestrate this with Kestra

Starting a process instance through Camunda's own console or API answers one moment in time; it does not decide a cadence, classify the outcome variable into a three-state risk level, gate a retry behind two independent confirmations, or notify anyone. Kestra supplies the schedule, the HEALTHY/AT_RISK/UNREACHABLE classification as a first-class branch, an authorization gate pair before any retry instance is started, and an execution history that shows exactly when a process stopped reaching its expected outcome.

Prerequisites

  • A reachable Camunda 8 cluster (self-managed or SaaS) with process_id already deployed, modeled to set expected_outcome_variable before every end event it can reach.
  • Credentials with permission to create process instances - this flow uses basic auth (username/password); OAuth2 (clientId/clientSecret/authorizationServerUrl/audience) is also supported by the base connection if your cluster requires it instead.
  • A Slack incoming webhook for alerts.

Local testing: Camunda 8 self-managed is a multi-component distributed system (Zeebe gateway, Elasticsearch or OpenSearch, Operate, Tasklist, and usually an identity/auth component) - there is no single official docker run one-liner equivalent to the other blueprints in this repo. Use Camunda's own official "Camunda 8 Run" local quickstart distribution or self-managed docker-compose.yaml (see the Camunda docs linked below) rather than a hand-assembled set of containers; whichever you choose, confirm its REST API port before setting camunda_rest_address, and avoid host port 8080 if Kestra's own UI is already bound there - map Camunda's REST gateway to a different host port instead.

Secrets

  • CAMUNDA_USERNAME / CAMUNDA_PASSWORD: basic-auth credentials used by every io.kestra.plugin.camunda.CreateProcessInstance task in this flow.
  • SLACK_WEBHOOK_URL: Slack incoming webhook used by alert_outcome_at_risk, alert_outcome_unreachable, and the errors block.
  • In Kestra OSS (no Enterprise secrets backend), secrets are supplied as environment variables prefixed SECRET_, base64-encoded, and read back in flows with {{ secret('NAME') }} - for example SECRET_CAMUNDA_PASSWORD=$(echo -n '<your-local-password>' | base64). This keeps credentials out of the flow YAML but, per Kestra's own documentation, offers no encryption at rest or access control beyond the host environment; use the Enterprise secrets backend for stronger guarantees.

Inputs

  • camunda_rest_address (STRING, default http://localhost:8080): Camunda REST API base URL.
  • process_id (STRING, default order-fulfillment): BPMN process id started and watched.
  • expected_outcome_variable (STRING, default status): variable name the model is expected to set.
  • expected_outcome_value (STRING, default approved): value that variable must equal for HEALTHY.
  • auto_remediate (SELECT: "false", "true"; default "false"): must be "true" for a retry to even be considered.
  • dry_run (SELECT: "true", "false"; default "true"): must be explicitly "false", together with auto_remediate: "true", for a retry instance to actually start.

Outputs

  • outputs.run_monitored_process.processInstanceKey / .variables: the monitored instance's identity and fetched outcome variable, every run.
  • outputs.retry_process_instance.processInstanceKey / .variables: present only when a retry actually ran.

Quick start

  1. Deploy a BPMN process under process_id that sets a status variable to approved on its normal end event (and something else on any rejection path) to your Camunda cluster.
  2. Add the CAMUNDA_USERNAME, CAMUNDA_PASSWORD, and SLACK_WEBHOOK_URL secrets.
  3. Run the flow manually with the defaults and confirm log_outcome_healthy fires on a normal run that reaches the approved path.
  4. Drive a test instance down a rejection path (or temporarily remodel the end event) and re-run to confirm alert_outcome_at_risk fires.
  5. Run once with auto_remediate: "true" and dry_run: "true" and confirm log_dry_run_remediation describes the retry without starting it.
  6. Run once with auto_remediate: "true" and dry_run: "false" and confirm retry_process_instance starts a new instance, visible in Camunda Operate.
  7. Point expected_outcome_variable at a variable name the model never sets and re-run to confirm alert_outcome_unreachable fires (UNREACHABLE) instead of either HEALTHY or AT_RISK.
  8. Enable periodic_outcome_check once you trust the check; it always runs in the safe auto_remediate: "false" / dry_run: "true" mode regardless of what you leave the flow's own defaults set to.

How to extend

  • Loop over several process_id/expected_outcome_variable pairs with io.kestra.plugin.core.flow.Loop to run the same outcome gate across every critical process in one flow.
  • Pass variables on run_monitored_process with real business input (an order ID, a customer tier) instead of starting the process with no input data, if the modeled process requires it to reach a meaningful outcome.
  • Add a maximum retry count read from the KV store before retry_process_instance, so a persistently rejecting process does not get retriggered indefinitely on every scheduled run.
  • Add io.kestra.plugin.camunda.Job-based handling for incidents that stop a process entirely, as a case this flow does not cover - an incident is a different failure mode than a completed instance with an unwanted outcome.

Pitfalls

  • This task has no success/failure state field, only whatever variables the BPMN model sets. Unlike io.kestra.plugin.airflow.dags.TriggerDagRun, which this repo's airflow-dag-run-sla-gate.yaml reads a state field from, CreateProcessInstance returns only variables on awaitCompletion - the entire classification here depends on the modeled process actually recording an outcome, which is a modeling discipline this flow cannot enforce.
  • This flow triggers real new process instances every time it runs, including the scheduled trigger. It is not a passive status check - run_monitored_process starts real work on every execution. Do not point process_id at a process where an extra, unplanned instance would be unsafe or expensive.
  • An incident (a Camunda-level execution failure inside the process) is not the same as an AT_RISK outcome. An instance stuck on an incident never reaches an end event, so awaitCompletion keeps waiting until requestTimeout and then this task throws, caught by the errors block - it is never classified as AT_RISK or UNREACHABLE by the branches in this flow.
  • requestTimeout must exceed how long the process actually takes. Per the plugin's own documentation, the client default of 10 seconds is usually too short when awaitCompletion is enabled; this flow sets PT5M, but a genuinely slow process needs a longer value on both run_monitored_process and retry_process_instance.
  • The retry only ever starts a new instance; it never touches the original one. retry_process_instance does not cancel, inspect, or resolve the first AT_RISK/UNREACHABLE instance - it is a fresh attempt, not a repair of the one just observed.
  • auto_remediate and dry_run are independent gates, not a single boolean. Both must be "true"/"false" respectively for a retry to actually start; setting only one leaves the other still blocking.
  • Switch case keys "true"/"false" are quoted strings. auto_remediate renders as the literal string "true" or "false"; unquoted true:/false: YAML map keys would parse as booleans instead and would not match.
  • This blueprint is UNTESTED against a live Camunda cluster. CreateProcessInstance's properties and outputs, and AbstractCamundaTask's connection fields, are taken verbatim from the plugin's source on main - but no Docker/Camunda environment or Kestra engine was run to execute this flow end to end, and local Camunda 8 setup was intentionally not reduced to a single invented docker run command.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.