Queries icon
Query icon
If icon
Loop icon
Set icon
SlackIncomingWebhook icon
Create icon
Log icon
Webhook icon

LLM Token and Dollar Budget Guardrail for Teams

Stop runaway LLM bills in Kestra: gateway usage events land in Postgres, over-budget teams get locked, alerted in Slack, and filed as GitHub issues.

Categories
AI

LLM spend rarely explodes in one call - it explodes one small charge at a time, across hundreds of calls, until the invoice arrives. A team wires up a new agent, another leaves a loop running, and nobody notices because per-call costs look trivial in isolation. This blueprint turns budget enforcement into an event-driven workflow: the gateway posts every usage record, a single query compares today's spend against per-team caps, and the moment a team crosses its limit it is locked, notified, and tracked to closure.

How it works

  1. record_usage (io.kestra.plugin.jdbc.postgresql.Queries, fetchType: NONE) receives the usage event from your LLM gateway webhook - team, model, prompt and completion tokens, and the cost of that call - and appends it to the llm_usage table.
  2. daily_spend_check (io.kestra.plugin.jdbc.postgresql.Query, FETCH) renders your budgets input as SQL VALUES, aggregates today's spend per team (all models) or per team-and-model (specific model caps), and returns only the rows at or over their limit.
  3. budget_gate (io.kestra.plugin.core.flow.If) branches on the breach. Inside a Loop over the offending teams, set_budget_lock (io.kestra.plugin.core.kv.Set) writes a llm_budget_lock/<team> KV key that a gateway sidecar or a parent flow can read to refuse new calls, and alert_team posts the exact numbers to Slack - spent, cap, tokens.
  4. open_issue (io.kestra.plugin.github.issues.Create) opens one labeled issue with a table of every offending team, so budget breaches get worked like any other engineering ticket instead of vanishing into a chat channel.
  5. The flow-level errors handler alerts Slack when the guard itself fails, because unmeasured spend is worse than measured spend.

What you get

  • Per-team daily dollar caps with optional per-model precision - one JSON input, no SQL edits.
  • A durable usage ledger in Postgres you can slice by team, model, or day.
  • Automatic quarantine keys (llm_budget_lock/<team>) for downstream enforcement.
  • over_budget_count and the full offender list as flow outputs for dashboards.
  • A GitHub issue trail - budget breaches become assignable, closable work.

Who it's for

  • Platform teams running shared LLM gateways (LiteLLM, OpenRouter, custom proxies) across many product teams.
  • FinOps owners who need per-team chargeback and hard stops instead of retrospective invoices.
  • Engineering leads who want an agent experiment to pause at $50, not surface at $5,000.

Why orchestrate this with Kestra

Budget logic embedded in a gateway dies when the gateway changes, and a cron script that checks costs alerts nobody at 2 a.m. Kestra gives you a webhook trigger on every usage event, a failure path that pages instead of hiding, KV state for the lock, and a loop that fans out alerts per team - all visible in an execution history you can audit when a team disputes its bill.

Prerequisites

  • PostgreSQL 12 or reachable from Kestra with the usage table:

    CREATE TABLE llm_usage (
      id BIGSERIAL PRIMARY KEY,
      team TEXT NOT NULL,
      model TEXT NOT NULL,
      prompt_tokens BIGINT NOT NULL DEFAULT 0,
      completion_tokens BIGINT NOT NULL DEFAULT 0,
      cost_usd NUMERIC(12,6) NOT NULL,
      occurred_at TIMESTAMPTZ NOT NULL DEFAULT now()
    );
    
  • An LLM gateway, proxy, or SDK wrapper that can POST a JSON usage record per completed call.

  • A Kestra instance.

Secrets

  • POSTGRES_JDBC_URL: JDBC URL for the usage database (for example jdbc:postgresql://host:5432/llmops).
  • POSTGRES_USERNAME: database username.
  • POSTGRES_PASSWORD: database password.
  • SLACK_WEBHOOK_URL: Slack incoming webhook for breach and failure alerts.
  • GITHUB_TOKEN: token with issue-write access to the repository input.

Quick start

  1. Add the five secrets above and create the llm_usage table.
  2. Set the budgets input to your real teams and caps, and point repository at your ops repo.
  3. Point your gateway's usage/cost webhook at the usage_events trigger URL.
  4. Fire a test event ({"team":"research","model":"gpt-4o","prompt_tokens":1200,"completion_tokens":400,"cost_usd":0.02}) and confirm the row lands.
  5. Lower a cap below your test spend and fire again - expect the KV lock, the Slack alert, and the GitHub issue.

How to extend

  • Enforce the lock: have your gateway read llm_budget_lock/<team> from KV and reject or queue calls while the key exists (keys are naturally reset daily - clear them with a scheduled flow at midnight).
  • Add a disabled daily schedule and gate record_usage behind an If on trigger.body, so a sweep re-checks spend even when a gateway batches its webhooks.
  • Escalate repeated breaches: chain a PagerDuty or Squadcast task after alert_team for teams that breach on consecutive days.
  • Feed chargeback: copy llm_usage daily aggregates into your billing warehouse with a Flow trigger chained from this one.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.