Schedule icon
Loop icon
ChatCompletion icon
OpenAI icon
Aggregate icon
OutputValues icon
Log icon
Fail icon
SlackIncomingWebhook icon

LLM Evaluation & Regression Testing

Evaluate LLM responses against a golden dataset, calculate correctness, relevance, and groundedness scores, and detect regressions with Kestra.

Categories
AI

Build a reusable evaluation pipeline for Large Language Model applications. This blueprint runs a model against a configurable golden dataset, evaluates every response using LLM-as-a-Judge, aggregates quality metrics, and detects regressions against a configurable baseline.

How it works

  1. The daily_evaluation schedule trigger is disabled by default and can be enabled for recurring evaluations.
  2. The evaluate_dataset loop processes every item in the golden dataset with configurable concurrency.
  3. The generate_answer task sends each question to the configured LLM through the OpenAI-compatible Groq endpoint.
  4. The judge task evaluates each response for correctness, relevance, and groundedness using structured JSON output.
  5. The aggregate_scores task calculates average scores across the complete evaluation dataset.
  6. The calculate_quality task calculates a single overall quality score from the three evaluation dimensions.
  7. The quality_threshold_failure task fails the execution when the overall score is below the required quality threshold.
  8. The regression_failure task detects quality degradation relative to the configured baseline and tolerance.
  9. Optional Slack notifications can be enabled for successful and failed evaluations.

What you get

  • Per-test evaluation results.
  • Correctness, relevance, and groundedness metrics.
  • A single overall quality score.
  • Configurable minimum quality threshold.
  • Configurable regression baseline and tolerance.
  • Parallelized dataset evaluation through Kestra's Loop task.
  • Structured JSON evaluation results.
  • Optional Slack notifications.
  • A reusable evaluation workflow that can be adapted to different LLM applications.

Who it's for

  • AI and ML engineers testing LLM applications.
  • Teams maintaining production RAG or GenAI applications.
  • Developers introducing automated LLM regression testing.
  • Platform teams building repeatable model evaluation pipelines.

Why orchestrate LLM evaluation with Kestra

LLM applications can change behavior when models, prompts, providers, or application code change. Manual testing does not scale well for repeated evaluations.

Kestra provides declarative orchestration, concurrency, scheduling, structured outputs, execution history, failure handling, and reusable workflow configuration in a single YAML definition.

Prerequisites

  • A Kestra instance.
  • An OpenAI-compatible LLM provider.
  • A Groq API key for the default configuration.
  • A GROQ_API_KEY Kestra secret.
  • Optional SLACK_WEBHOOK secret if Slack notifications are enabled.

Quick start

  1. Configure the GROQ_API_KEY secret in your Kestra namespace.
  2. Review the golden_dataset input and replace the examples with your own evaluation cases.
  3. Adjust model_name, quality_threshold, baseline_score, and regression_tolerance as required.
  4. Execute the flow manually to inspect individual and aggregate evaluation results.
  5. Enable the daily_evaluation trigger when recurring evaluation is required.
  6. Configure SLACK_WEBHOOK and remove disabled: true from the Slack notification task and error handler if Slack alerts are desired.

How to extend

  • Replace the default OpenAI-compatible provider with another supported Kestra AI provider.
  • Add additional evaluation dimensions such as factuality, safety, or task success.
  • Add deterministic code-based evaluators for metrics that do not require an LLM judge.
  • Store historical evaluation scores in a database for long-term regression tracking.
  • Add model comparison by evaluating multiple model inputs against the same golden dataset.
  • Add downstream reporting or dashboard tasks using the generated evaluation outputs.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.