Schedule icon
Webhook icon
Script icon
Process icon
If icon
SlackIncomingWebhook icon
Log icon

Evaluate RAG Pipelines for Hallucinations and Faithfulness

Evaluate RAG pipelines for hallucinations and ungrounded claims using an LLM-as-a-judge, generate audit reports, and alert Slack on regressions.

Categories
AICloud

In production Retrieval-Augmented Generation (RAG) applications, subtle revisions to system prompts, embedding models, or document chunking strategies can quietly introduce hallucinations. Generative drift can cause models to invent facts, dates, features, or policy terms that are completely unsupported by the retrieved document chunks.

This blueprint implements an automated, repeatable RAG quality gate powered by an LLM-as-a-judge. It systematically scores generated answers against retrieved document context, detects ungrounded assertions and factual fabrications, saves an audit report artifact, and triggers an immediate Slack notification when faithfulness falls below your tolerance threshold.

How it works

  1. evaluate_rag_faithfulness (io.kestra.plugin.scripts.python.Script on the io.kestra.plugin.core.runner.Process runner) iterates across your evaluation suite, sending query, retrieved context, and generated answer triples to an LLM judge. The judge performs sentence-level claim decomposition, assigns a normalized groundedness score (0.0 to 1.0), and flags any claims absent from the retrieved chunks.
  2. generate_audit_report (io.kestra.plugin.scripts.python.Script) parses the evaluation metrics and compiles a downloadable Markdown compliance report (rag-hallucination-audit-report.md) persisted into Kestra internal storage.
  3. evaluate_faithfulness_gate (io.kestra.plugin.core.flow.If) assesses whether hallucination_rate exceeds max_hallucination_rate or if avg_faithfulness drops below min_faithfulness_score. If a regression is detected, alert_hallucination_breach dispatches an alert with failing query IDs to Slack; otherwise, log_clean_audit records a passing audit.
  4. The errors block alerts Slack if the evaluation harness encounters API connection errors or authentication failures.
  5. Triggers: a nightly Schedule (shipped disabled) to continuously catch quiet model drift, and an event-driven Webhook to execute the gate automatically in CI/CD pipelines on prompt or index updates.

What you get

  • Automated LLM-as-a-judge groundedness evaluation for production RAG pipelines.
  • Zero external Python package dependencies (runs purely on Python standard library).
  • Detection and extraction of specific hallucinated claims per test case.
  • Downloadable Markdown audit artifact (rag-hallucination-audit-report.md) with executive summaries and actionable tuning recommendations.
  • Machine-readable eval_summary JSON output with mean scores, breach rates, and failing query IDs.
  • Real-time Slack notifications for rapid ML engineering and security triage.

Who it's for

  • AI Engineers deploying RAG systems who need automated regression testing before production releases.
  • MLOps and DevSecOps teams establishing automated quality and compliance gates for GenAI pipelines.
  • Platform engineers managing vector search indexes, knowledge bases, and LLM endpoints.

Why orchestrate this with Kestra

Evaluating RAG quality cannot be a static, one-time exercise. Documents, chunking parameters, embedding indexes, and model weights evolve continuously. Kestra orchestrates scheduled automated audits, CI/CD webhook triggers, failure handling, artifact persistence, and Slack alerting in a declarative workflow, while recording execution trends across model versions over time.

Prerequisites

  • An OpenAI-compatible chat completions endpoint (OpenAI, Ollama, vLLM, or LocalAI).
  • A Slack incoming webhook URL for alert notifications.

Secrets

  • OPENAI_API_KEY: API key for the LLM judge model.
  • SLACK_WEBHOOK_URL: Slack webhook URL for hallucination regression alerts.

Quick start

  1. Configure the OPENAI_API_KEY and SLACK_WEBHOOK_URL secrets in your Kestra namespace.
  2. Import this blueprint and customize the rag_eval_suite JSON input with queries and context chunks from your domain.
  3. Execute the flow manually to inspect the generated rag-hallucination-audit-report.md artifact.
  4. Enable the nightly_rag_audit schedule or trigger via audit_webhook in your CI/CD pipeline.

How to extend

  • Connect to a vector database plugin (such as Pinecone, Qdrant, or Weaviate) to dynamically fetch live context chunks.
  • Add automated rollback subflows to revert model endpoints or prompt revisions when regressions occur.
  • Add custom evaluation criteria such as answer relevance, context precision, and toxicity scoring.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.