Schedule icon
Webhook icon
Log icon
Loop icon
ChatCompletion icon
OpenAI icon
OutputValues icon
Write icon
If icon
SlackIncomingWebhook icon
Fail icon

RAG Synthetic Golden Dataset Generator for LLM Evaluation

Generate synthetic Q&A golden datasets from document chunks using Kestra's AI plugin with structured JSON schema for automated RAG and LLM evaluation.

Categories
AIData

Evaluating Large Language Model (LLM) and Retrieval-Augmented Generation (RAG) pipelines requires a robust "Golden Dataset" composed of realistic questions, ground-truth answers, and reference context chunks. Manually authoring hundreds of domain-specific evaluation pairs is cost-prohibitive, tedious, and difficult to keep in sync with evolving knowledge bases.

This blueprint automates synthetic golden dataset creation by orchestrating LLM-as-a-Generator across raw document chunks using Kestra's native AI plugin with strict JSON Schema outputs, verifying test case coverage, and exporting a versioned golden dataset artifact.

How it works

  1. log_generation_start (io.kestra.plugin.core.log.Log) logs the synthesis run parameters and total input chunks.
  2. generate_synthetic_cases (io.kestra.plugin.core.flow.Loop) iterates through each document chunk in parallel with configurable concurrency.
  3. synthesize_qa (io.kestra.plugin.ai.completion.ChatCompletion) sends each chunk to an OpenAI-compatible LLM endpoint with a strict JSON Schema, enforcing extraction of question, ground-truth expected answer, verbatim context excerpt, difficulty rating, and technical topic.
  4. aggregate_dataset (io.kestra.plugin.core.output.OutputValues) aggregates synthesized test cases across all loop outputs into a single flattened collection.
  5. calculate_metrics (io.kestra.plugin.core.output.OutputValues) measures total generated evaluation cases.
  6. export_golden_dataset (io.kestra.plugin.core.storage.Write) persists the consolidated golden dataset into Kestra internal storage as a downloadable JSON artifact.
  7. dataset_quality_gate (io.kestra.plugin.core.flow.If) verifies that the generated test count meets or exceeds min_total_cases. If successful, it logs the artifact URI and sends an optional Slack notification; otherwise, it triggers fail_insufficient_cases (io.kestra.plugin.core.execution.Fail).
  8. The top-level errors handler captures runtime failures and sends an immediate alert to Slack.

What you get

  • Automated generation of domain-specific Q&A evaluation datasets from raw documentation.
  • Strict JSON Schema validation guaranteeing test case structure (id, question, expected, context, difficulty, topic).
  • Parallel chunk processing with controlled concurrency using Kestra's native Loop.
  • Golden dataset persisted directly to Kestra internal storage as a versioned JSON artifact.
  • Minimum coverage quality gate preventing downstream evaluation on incomplete datasets.
  • Seamless compatibility with downstream evaluation blueprints such as llm-evaluation-regression.
  • Proactive Slack notifications for successful generation and execution errors.

Who it's for

  • AI and ML Engineers building and evaluating production RAG pipelines.
  • Platform teams creating continuous evaluation CI/CD loops for knowledge bases.
  • Data Engineers responsible for keeping evaluation datasets aligned with new documentation releases.
  • Teams adopting LLM-as-a-Judge workflows who need automated ground-truth generation.

Why orchestrate this with Kestra

Generating synthetic datasets over large document corpora involves concurrency control, rate-limit handling, structured output validation, state persistence, and error alerting. Writing ad-hoc Python scripts leaves teams with untracked files, unhandled API timeouts, and no audit trail.

Kestra provides declarative orchestration, native concurrency throttling, built-in internal storage management, execution history, and seamless integration between AI generation tasks and notification systems—all defined in readable YAML.

Prerequisites

  • A running Kestra instance.
  • An OpenAI-compatible LLM API key (Groq, OpenAI, or self-hosted vLLM).
  • An optional Slack incoming webhook URL for notifications.

Secrets

  • GROQ_API_KEY: API key for accessing the LLM endpoint (default uses Groq's OpenAI-compatible API).
  • SLACK_WEBHOOK_URL: Webhook URL for Slack alerts (optional, enabled by removing disabled: true).
  • WEBHOOK_KEY: Secret authentication key for the inbound webhook trigger.

Quick start

  1. Configure GROQ_API_KEY (and optionally SLACK_WEBHOOK_URL) in your Kestra namespace.
  2. Provide your document chunks via the document_chunks input, or connect an upstream ingestion task.
  3. Run the workflow manually from the Kestra UI to verify synthesis output.
  4. Download the generated golden_dataset.json artifact from the execution outputs, or pass its URI directly into llm-evaluation-regression.
  5. Enable weekly_dataset_refresh or wire your documentation build pipeline to documentation_update_webhook.

How to extend

  • Ingest document chunks directly from an AWS S3 bucket, Git repository, or vector database before generation.
  • Feed the resulting export_golden_dataset.uri directly into the llm-evaluation-regression blueprint via a subflow execution.
  • Add multi-language synthesis prompts to generate multilingual RAG evaluation sets.
  • Incorporate hard negative generation (synthesizing unanswerable questions to test hallucination resistance).

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.