Schedule icon
Webhook icon
Script icon
Process icon
If icon
SlackIncomingWebhook icon
Log icon

Gate Every Prompt and Model Change Behind a Golden Eval Set

Run a golden set of prompts through your model, grade each answer, and alert Slack the moment the pass rate falls below your release bar.

Categories
AICore

Prompt tweaks and model swaps rarely break loudly — they shave a few percent off answer quality and nobody notices until support does. This blueprint makes that impossible to miss: a fixed golden set of prompts runs on a schedule or from CI, every answer is graded against its expected marker, and the pass rate is compared with a threshold before the change ships. Failing case ids land in Slack, passing scores land in the execution history, and the trend across versions tells you whether the new model is actually better.

How it works

  1. run_eval (io.kestra.plugin.scripts.python.Script on the io.kestra.plugin.core.runner.Process runner) loads the golden cases from the eval_cases input, calls the OpenAI-compatible /chat/completions endpoint once per case with temperature: 0 using only the Python standard library, and grades each reply by case-insensitive substring match against expected_contains. Transient HTTP errors get one retry; 401/403 fail fast. Totals, pass rate, and failing ids are emitted through the ::{"outputs": ...}:: protocol.
  2. check_pass_rate (io.kestra.plugin.core.flow.If) compares pass_rate with the pass_rate_threshold input. Below it, alert_regression posts the score, the bar, the model, and the failing case ids to Slack; at or above it, log_pass writes the score to the execution history so runs become a version-over-version trend.
  3. The errors block alerts Slack when the harness itself breaks — a rejected key or dead endpoint must never read as a passing score.
  4. Triggers: a nightly Schedule (shipped disabled) plus a Webhook so CI can call the gate on the exact commit that changed a prompt or model id.

What you get

  • A deterministic release gate: model or prompt changed → pass rate measured → pass or alert, same cases every time.
  • Failing case ids in the alert — the exact prompts to reopen, not just a number.
  • An eval_summary JSON output (total, passed, pass_rate, failed_ids) for dashboards or a downstream approval flow.
  • A score history in the execution log, one entry per run.

Who it's for

  • Teams shipping LLM features who want a regression test for prompts the way unit tests cover code.
  • Platform engineers comparing a new model against the old one before switching traffic.
  • Anyone who has promoted a prompt change on vibes and found out a week later.

Why orchestrate this with Kestra

A test script can score a golden set, but it cannot run itself nightly, fire a webhook from CI, alert the right channel on failure, self-report harness errors, and keep every score in a queryable history. Kestra adds the schedule and webhook triggers, the threshold branch, the failure channel, and — because the gate is a flow — the natural place to later add an approval task or a compare-with-yesterday step.

Prerequisites

  • An OpenAI-compatible endpoint and API key (OpenAI, or any compatible server — set base_url to point at Ollama, vLLM, etc.).
  • A Slack incoming webhook for regression and harness-failure alerts.
  • Golden cases with honest expected markers: substrings that a good answer must contain. Keep them stable — the value of the gate is that the set never changes silently.

Secrets

  • OPENAI_API_KEY: bearer token sent to the chat completions endpoint.
  • SLACK_WEBHOOK_URL: Slack incoming webhook URL for regression and failure alerts.

Quick start

  1. Add the two secrets to your Kestra namespace.
  2. Replace eval_cases with your own prompts and expected markers.
  3. Run the flow once and read eval_summary to confirm the score matches what you expect.
  4. Enable the nightly_eval schedule, and POST to the llm-eval-gate webhook from CI when prompts change.

How to extend

  • Swap the substring grader for a judge model: add a second io.kestra.plugin.openai.ChatCompletion per case that scores the answer against the rubric and aggregates the judge's verdicts.
  • Store eval_summary in KV and alert only when the rate drops relative to yesterday, not just below the absolute bar.
  • Split the gate by case group (safety, reasoning, formatting) with one threshold per group by adding an If per group.
  • Fan a failure out to a Jira or Linear ticket by adding a second task in the then branch.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.