Webhook icon
Schedule icon
Request icon
If icon
Fail icon
WorkingDirectory icon
Script icon
Process icon
IngestDocument icon
Ollama icon
KestraKVStore icon
Loop icon
Parallel icon
Search icon
Queries icon
SlackIncomingWebhook icon
Exit icon
Log icon

Check What Changes Before You Switch Embedding Models

Compare current and candidate embedding models on your own documents and probe queries, score top-k overlap, and block risky migrations.

Categories
AIData

Switching embedding models is one of the riskiest changes in a RAG system: every stored vector is re-computed, and the documents your users get back change in ways no unit test sees. This blueprint measures that change before you ship it. It indexes the same documents with the current and the candidate model through the Kestra AI plugin, runs a fixed set of probe queries against both indexes, and compares the top-k results. When the average overlap falls below your threshold, the migration is blocked and Slack lists the queries that changed the most.

It runs fully locally with Ollama by default, so you can try it without API keys or cost, and the provider blocks can be swapped for OpenAI, Gemini, Mistral, or any other AI plugin provider.

How it works

  1. list_models + check_models (If) confirm both models are installed in Ollama. If one is missing, the run fails within seconds with the exact ollama pull command to run.
  2. build_indexes (WorkingDirectory) runs three tasks on the same files:
    • chunk_documents downloads document_urls and writes one file per chunk, each prefixed with its id ([c0001] ...)
    • index_current and index_candidate (io.kestra.plugin.ai.rag.IngestDocument) embed the identical chunks with each model into separate KestraKVStore indexes
  3. probe (Loop) runs each probe query against both indexes in parallel with io.kestra.plugin.ai.rag.Search.
  4. compare (io.kestra.plugin.jdbc.duckdb.Queries) matches chunk ids between the two result lists and scores each query:
    • overlap: share of the top-k returned by both models
    • top1_same: whether the best chunk is unchanged
  5. gate (If) blocks the migration when the average overlap is below min_overlap: Slack gets the score and the three most-changed queries, and the execution ends FAILED so a CI job polling it fails too. Otherwise the result is logged.
  6. The errors block alerts when the check itself cannot run.

Error handling

The flow separates two outcomes that a CI job must not confuse.

The check ran and the migration is blocked.

  • blocked uses Exit with state FAILED, which ends the run without triggering the flow's errors block.
  • Slack gets one "BLOCKED" message with the score and the most-changed queries.

The check could not run. These failures use Fail or a task error, so the errors block posts the failing task and the cause:

  • A missing model fails in check_models before any indexing, for example "Embedding model missing on http://localhost:11434. Install it with: ollama pull mxbai-embed-large;".
  • An unreachable Ollama fails in list_models after three attempts with "Connection refused".
  • Transient errors from the embedding server are retried: every AI task retries with exponential backoff.
  • An empty corpus stops chunk_documents with "no text found in document_urls".

What this blueprint teaches about Kestra

  • The AI plugin's RAG tasks (IngestDocument, Search) used as a measurement tool, not only for chat: two embedding stores side by side, compared in DuckDB.
  • Loop combined with Parallel, and loop outputs collected into a single JSON file for SQL analysis.
  • Exit versus Fail: one ends a run with a deliberate verdict, the other is an error that triggers alerting.

What you get

  • summary: average top-k overlap, top-1 agreement, and the number of queries below the threshold.
  • per_query: for every probe query, the overlap and the ranked chunk ids from each model, most changed first, so you know exactly which questions to re-test.
  • A go/no-go gate you can call from CI through the webhook on any pull request that changes the embedding model.

Prerequisites

  • An Ollama server reachable from the worker, with both models pulled (ollama pull all-minilm, ollama pull nomic-embed-text), or credentials for another AI plugin provider.
  • Python 3 on the worker for chunk_documents (standard library only, Process task runner).
  • A Slack incoming webhook.

Secrets

  • SLACK_WEBHOOK_URL: Slack incoming webhook URL for blocked migrations and failures.

Inputs

  • document_urls (ARRAY): documents that make up the test corpus. Defaults to the Kestra README.
  • probe_queries (ARRAY): questions your users ask; defaults to ten questions about Kestra.
  • current_model, candidate_model (STRING): defaults all-minilm and nomic-embed-text.
  • ollama_endpoint (STRING, default http://localhost:11434).
  • top_k (INT, default 5), min_overlap (FLOAT, default 0.6), chunk_chars (INT, default 600).

Outputs

  • summary (JSON) and per_query (JSON), as described above. They are set on passing runs; on a blocked run the same numbers are in the Slack alert and the compare task outputs.

Quick start

  1. Start Ollama and pull both models.
  2. Add the SLACK_WEBHOOK_URL secret.
  3. Run the flow with the defaults. With the Kestra README, all-minilm and nomic-embed-text share 54% of their top-5 results and agree on the best chunk for half of the queries, so the default 0.6 threshold blocks the switch.
  4. Replace document_urls and probe_queries with your own corpus and questions, and call the embedding-migration-check webhook from CI.

How to read the result

Low overlap does not mean the candidate is worse; it means retrieval changes. Use per_query to review the changed queries: if the candidate's chunks answer them better, lower the threshold or update your expectations, then migrate. Comparing a model with itself returns an overlap of 1.0, which is a quick way to confirm the setup.

How to extend

  • Compare hosted models: replace the Ollama provider with OpenAI or GoogleGemini in the four AI tasks and add the API key secret.
  • Score answer quality, not only retrieval: add a ChatCompletion task per query that answers from each model's chunks and a judge that picks the better answer.
  • Use a real vector store: swap KestraKVStore for PGVector, Qdrant, or Weaviate so large corpora fit.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.