New to Kestra?
Use blueprints to kickstart your first workflows.
Generate synthetic Q&A golden datasets from document chunks using Kestra's AI plugin with structured JSON schema for automated RAG and LLM evaluation.
Evaluating Large Language Model (LLM) and Retrieval-Augmented Generation (RAG) pipelines requires a robust "Golden Dataset" composed of realistic questions, ground-truth answers, and reference context chunks. Manually authoring hundreds of domain-specific evaluation pairs is cost-prohibitive, tedious, and difficult to keep in sync with evolving knowledge bases.
This blueprint automates synthetic golden dataset creation by orchestrating LLM-as-a-Generator across raw document chunks using Kestra's native AI plugin with strict JSON Schema outputs, verifying test case coverage, and exporting a versioned golden dataset artifact.
log_generation_start (io.kestra.plugin.core.log.Log) logs the synthesis run parameters and total input chunks.generate_synthetic_cases (io.kestra.plugin.core.flow.Loop) iterates through each document chunk in parallel with configurable concurrency.synthesize_qa (io.kestra.plugin.ai.completion.ChatCompletion) sends each chunk to an OpenAI-compatible LLM endpoint with a strict JSON Schema, enforcing extraction of question, ground-truth expected answer, verbatim context excerpt, difficulty rating, and technical topic.aggregate_dataset (io.kestra.plugin.core.output.OutputValues) aggregates synthesized test cases across all loop outputs into a single flattened collection.calculate_metrics (io.kestra.plugin.core.output.OutputValues) measures total generated evaluation cases.export_golden_dataset (io.kestra.plugin.core.storage.Write) persists the consolidated golden dataset into Kestra internal storage as a downloadable JSON artifact.dataset_quality_gate (io.kestra.plugin.core.flow.If) verifies that the generated test count meets or exceeds min_total_cases. If successful, it logs the artifact URI and sends an optional Slack notification; otherwise, it triggers fail_insufficient_cases (io.kestra.plugin.core.execution.Fail).errors handler captures runtime failures and sends an immediate alert to Slack.id, question, expected, context, difficulty, topic).Loop.llm-evaluation-regression.Generating synthetic datasets over large document corpora involves concurrency control, rate-limit handling, structured output validation, state persistence, and error alerting. Writing ad-hoc Python scripts leaves teams with untracked files, unhandled API timeouts, and no audit trail.
Kestra provides declarative orchestration, native concurrency throttling, built-in internal storage management, execution history, and seamless integration between AI generation tasks and notification systems—all defined in readable YAML.
GROQ_API_KEY: API key for accessing the LLM endpoint (default uses Groq's OpenAI-compatible API).SLACK_WEBHOOK_URL: Webhook URL for Slack alerts (optional, enabled by removing disabled: true).WEBHOOK_KEY: Secret authentication key for the inbound webhook trigger.GROQ_API_KEY (and optionally SLACK_WEBHOOK_URL) in your Kestra namespace.document_chunks input, or connect an upstream ingestion task.golden_dataset.json artifact from the execution outputs, or pass its URI directly into llm-evaluation-regression.weekly_dataset_refresh or wire your documentation build pipeline to documentation_update_webhook.export_golden_dataset.uri directly into the llm-evaluation-regression blueprint via a subflow execution.