Schedule icon
Webhook icon
Script icon
Process icon
If icon
Fail icon
Loop icon
WorkingDirectory icon
Commands icon
IngestDocument icon
OpenAI icon
KestraKVStore icon
Search icon
Queries icon
ChatCompletion icon
SlackIncomingWebhook icon
Log icon

Benchmark Conversational Memory Retrievers on LoCoMo with an LLM Judge

Compare BM25 and embedding retrieval on LoCoMo, score answers with F1 and an LLM judge, gate on cost, and alert Slack when recall drops.

Categories
AIData

A chatbot with long-term memory is only as good as what it retrieves from past conversations, but end-to-end QA scores hide where the points are lost. This blueprint separates the pieces on the public LoCoMo benchmark (long multi-session conversations with annotated evidence turns). Three retrievers answer the same question sample: lexical BM25, semantic embedding search, and an oracle that hands the reader the gold evidence. Retrieval is measured directly with recall@k, and every answer is scored twice, with token F1 and with an LLM judge. The gap between a retriever and the oracle tells you how much better retrieval would buy before you build it.

How it works

  1. prepare downloads the LoCoMo release, verifies its SHA-256 against the pinned hash, writes each conversation as one line per turn, and draws a deterministic sample of answerable questions (single-hop, multi-hop, temporal, open-domain). It also estimates the cost of the run.
  2. budget_gate (io.kestra.plugin.core.flow.If) fails the run before any model call when the estimate is above max_cost_usd.
  3. index_conversations (io.kestra.plugin.core.flow.Loop, one parallel branch per conversation) splits each conversation into one file per turn and embeds them with io.kestra.plugin.ai.rag.IngestDocument into a per-conversation KestraKVStore, so a search never crosses conversations.
  4. embedding_search (Loop) runs io.kestra.plugin.ai.rag.Search for every question against its own conversation's store.
  5. build_contexts (io.kestra.plugin.jdbc.duckdb.Queries) ranks turns with BM25 written in plain SQL (statistics per conversation, no extension needed), collects the embedding hits and the gold evidence, and builds one reader context per question and retriever, with recall@k against the evidence turns.
  6. answer_and_judge (Loop) calls io.kestra.plugin.ai.completion.ChatCompletion twice per question and retriever: the reader answers from the retrieved turns only, then the judge grades the answer CORRECT or WRONG. Both tasks use Kestra's built-in retry for transient API errors.
  7. leaderboard (DuckDB) computes SQuAD-style token F1 in SQL and returns three tables: one row per retriever, a per-category breakdown, and the actual token cost of the run.
  8. check_recall alerts Slack with the full leaderboard when the best non-oracle retriever is below min_recall; otherwise the leaderboard is written to the execution logs. The errors block reports any failed run to Slack.

What you get

  • A leaderboard per retriever: recall_at_k, token_f1, judge_accuracy, and f1_judge_disagreement (how often F1 and the judge disagree on the same answer, a direct read on how metric-dependent your score is).
  • A per-category breakdown, since temporal and multi-hop questions usually fail for different reasons than single-hop ones.
  • The actual cost of the run from the token counts the AI plugin reports, next to the pre-run estimate.
  • Every reader answer and judge verdict as its own task run in the Kestra UI, so a surprising score can be traced to the exact question.

Who it's for

  • Teams building chat assistants or agents with long-term memory who want to choose a retrieval strategy on evidence.
  • Researchers who need a repeatable, budgeted LoCoMo run they can trigger from CI.
  • Anyone who suspects their memory system's score depends more on the judge than on retrieval.

Prerequisites

  • An OpenAI-compatible API key with access to a chat model and an embedding model. Set base_url to use another compatible server.
  • A Slack incoming webhook for alerts.
  • Outbound access from the worker to raw.githubusercontent.com to download the dataset.
  • Python 3 and coreutils on the worker for prepare and split_turns, which run on the Process task runner. To isolate them in a container, swap the task runner for io.kestra.plugin.scripts.runner.docker.Docker.

Secrets

  • OPENAI_API_KEY: API key for the reader, judge, and embedding calls.
  • SLACK_WEBHOOK_URL: Slack incoming webhook URL for low-recall and failure alerts.

Inputs

  • dataset_url (URI) and dataset_sha256 (STRING): the LoCoMo release and its expected hash. Defaults to locomo10.json from the official repository, checked against its known hash.
  • max_conversations (INT, default 10), sample_size (INT, default 40), seed (STRING): which conversations to index and the deterministic question sample.
  • top_k (INT, default 5): turns each retriever passes to the reader.
  • model, judge_model, embedding_model, base_url (STRING): reader, judge, embedding model, and endpoint.
  • price_in_per_1m, price_out_per_1m, price_embedding_per_1m, max_cost_usd (FLOAT): pricing for the estimate and the cost report, and the budget for one run.
  • min_recall (FLOAT, default 0.5): recall@k bar for the best non-oracle retriever.

Outputs

  • leaderboard (JSON): one row per retriever, from {{ outputs.leaderboard.outputs[0].rows }}.
  • by_category (JSON): per-category rows, from {{ outputs.leaderboard.outputs[1].rows }}.
  • cost (JSON): chat and embedding token counts and actual_cost_usd, from {{ outputs.leaderboard.outputs[2].rows[0] }}.

Quick start

  1. Add the OPENAI_API_KEY and SLACK_WEBHOOK_URL secrets (in the open-source edition, as base64-encoded SECRET_OPENAI_API_KEY and SECRET_SLACK_WEBHOOK_URL environment variables).
  2. Run the flow once with the defaults. Forty questions across three retrievers on gpt-4o-mini costs about one cent.
  3. Open the leaderboard output and compare bm25 and embedding against oracle.
  4. Enable weekly_benchmark, or POST to the locomo-memory-benchmark webhook from CI when your retrieval code changes.

How to extend

  • Benchmark your own memory store: swap KestraKVStore for PGVector, Qdrant, or another embedding store from the AI plugin, or add a retriever to the hits table in build_contexts.
  • Keep a score history: append the leaderboard to a persistent DuckDB file or Postgres table and alert on a drop relative to the previous run instead of a fixed bar.
  • Calibrate the judge: export the answers where F1 and the judge disagree and label them by hand to measure how often the judge is right.

Notes

  • KestraKVStore keeps vectors in Kestra's KV store and loads them into memory on search. It is fine for LoCoMo-sized conversations; use a real vector store for production-sized memory.
  • Model outputs can vary slightly between runs even at temperature 0, so expect small run-to-run differences in F1 and judge accuracy for the same sample and settings. Retrieval and recall@k are deterministic for BM25 and the oracle.
  • Token F1 is the SQuAD-style metric without stemming, so absolute numbers differ slightly from the official LoCoMo scorer. Compare runs of this flow with each other, not with published leaderboards.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.