New to Kestra?
Use blueprints to kickstart your first workflows.
Stop runaway LLM bills in Kestra: gateway usage events land in Postgres, over-budget teams get locked, alerted in Slack, and filed as GitHub issues.
LLM spend rarely explodes in one call - it explodes one small charge at a time, across hundreds of calls, until the invoice arrives. A team wires up a new agent, another leaves a loop running, and nobody notices because per-call costs look trivial in isolation. This blueprint turns budget enforcement into an event-driven workflow: the gateway posts every usage record, a single query compares today's spend against per-team caps, and the moment a team crosses its limit it is locked, notified, and tracked to closure.
record_usage (io.kestra.plugin.jdbc.postgresql.Queries, fetchType: NONE) receives the usage event from your LLM gateway webhook - team, model, prompt and completion tokens, and the cost of that call - and appends it to the llm_usage table.daily_spend_check (io.kestra.plugin.jdbc.postgresql.Query, FETCH) renders your budgets input as SQL VALUES, aggregates today's spend per team (all models) or per team-and-model (specific model caps), and returns only the rows at or over their limit.budget_gate (io.kestra.plugin.core.flow.If) branches on the breach. Inside a Loop over the offending teams, set_budget_lock (io.kestra.plugin.core.kv.Set) writes a llm_budget_lock/<team> KV key that a gateway sidecar or a parent flow can read to refuse new calls, and alert_team posts the exact numbers to Slack - spent, cap, tokens.open_issue (io.kestra.plugin.github.issues.Create) opens one labeled issue with a table of every offending team, so budget breaches get worked like any other engineering ticket instead of vanishing into a chat channel.errors handler alerts Slack when the guard itself fails, because unmeasured spend is worse than measured spend.llm_budget_lock/<team>) for downstream enforcement.over_budget_count and the full offender list as flow outputs for dashboards.Budget logic embedded in a gateway dies when the gateway changes, and a cron script that checks costs alerts nobody at 2 a.m. Kestra gives you a webhook trigger on every usage event, a failure path that pages instead of hiding, KV state for the lock, and a loop that fans out alerts per team - all visible in an execution history you can audit when a team disputes its bill.
PostgreSQL 12 or reachable from Kestra with the usage table:
CREATE TABLE llm_usage (
id BIGSERIAL PRIMARY KEY,
team TEXT NOT NULL,
model TEXT NOT NULL,
prompt_tokens BIGINT NOT NULL DEFAULT 0,
completion_tokens BIGINT NOT NULL DEFAULT 0,
cost_usd NUMERIC(12,6) NOT NULL,
occurred_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
An LLM gateway, proxy, or SDK wrapper that can POST a JSON usage record per completed call.
A Kestra instance.
POSTGRES_JDBC_URL: JDBC URL for the usage database (for example jdbc:postgresql://host:5432/llmops).POSTGRES_USERNAME: database username.POSTGRES_PASSWORD: database password.SLACK_WEBHOOK_URL: Slack incoming webhook for breach and failure alerts.GITHUB_TOKEN: token with issue-write access to the repository input.llm_usage table.budgets input to your real teams and caps, and point repository at your ops repo.usage_events trigger URL.{"team":"research","model":"gpt-4o","prompt_tokens":1200,"completion_tokens":400,"cost_usd":0.02}) and confirm the row lands.llm_budget_lock/<team> from KV and reject or queue calls while the key exists (keys are naturally reset daily - clear them with a scheduled flow at midnight).record_usage behind an If on trigger.body, so a sweep re-checks spend even when a gateway batches its webhooks.alert_team for teams that breach on consecutive days.llm_usage daily aggregates into your billing warehouse with a Flow trigger chained from this one.