Schedule icon
Webhook icon
Script icon
Queries icon
If icon
Write icon
SlackIncomingWebhook icon
Log icon

RAG Knowledge Base Vector Reconciliation and Sync

Reconcile vector store drift with source knowledge base documents using cryptographic hashing, purge orphaned vectors, and cut LLM embedding costs in Kestra.

Categories
AICloudData

Diagram unavailable

We could not build the topology for this blueprint. The flow itself is valid, use the YAML on the left to run it.

When running AI-powered search across internal documentation, engineering runbooks, or customer support wikis, keeping vector embeddings synchronized with source files is an operational headache. Documentation changes constantly: security policies are amended, API reference guides update, and obsolete runbooks get deleted.

When vector databases (Pinecone, Qdrant, Weaviate, pgvector) are maintained via ad-hoc scripts or naive batch syncs, two costly problems emerge:

  1. Orphaned Hallucination Vectors: When a source document is deleted, its vector embeddings linger indefinitely in the vector index, leading internal search and support bots to cite deprecated runbooks or contradictory policies.
  2. Costly Full Re-Indexing: Teams frequently re-embed their entire documentation corpus from scratch on every commit, burning thousands of dollars in embedding API tokens and hitting provider rate limits.

This blueprint automates incremental drift reconciliation in Kestra: it scans source files, computes cryptographic SHA-256 content hashes, leverages DuckDB in-memory relational set differences to detect inserts, updates, and deletes, immediately purges orphaned vectors, and selectively re-embeds only modified chunks.

How it works

  1. scan_source_documents (io.kestra.plugin.scripts.python.Script) scans source markdown/text documents, generates metadata, and computes cryptographic SHA-256 hashes for each document.
  2. reconcile_vector_drift (io.kestra.plugin.jdbc.duckdb.Queries) compares the active source document manifest against the vector index tracking catalog in DuckDB in-memory SQL to identify to_insert, to_update, and to_delete sets.
  3. evaluate_drift_gate (io.kestra.plugin.core.flow.If) evaluates whether any vector drift was identified:
    • Drift Detected (true):
      • execute_pruning_and_sync (io.kestra.plugin.scripts.python.Script): Purges orphaned vector IDs from the vector store and chunks new/modified documents, tracking token efficiency metrics.
      • export_sync_audit_report (io.kestra.plugin.core.storage.Write): Generates an executive Markdown audit log (rag-sync-report.md) detailing all vector operations.
      • notify_slack_sync_complete (io.kestra.plugin.slack.notifications.SlackIncomingWebhook): Dispatches a structured summary card to the engineering Slack channel.
    • Clean State (false):
      • log_clean_state (io.kestra.plugin.core.log.Log): Records zero drift without consuming unnecessary vector database compute or LLM API calls.
  4. alert_reconciliation_failure (errors block) immediately alerts Slack if any reconciliation or purge step encounters an exception.

What you get

  • Zero orphaned vectors: obsolete document embeddings are purged immediately upon source deletion.
  • Up to 99% cost reduction in embedding API bills by eliminating full-corpus re-indexing.
  • Sub-millisecond set reconciliation powered by embedded DuckDB.
  • Publication-grade audit trail generated for AI safety, compliance, and governance audits.
  • Real-time Slack notifications alerting engineering teams to knowledge base drift.

Who it's for

  • AI Engineers and LLM Application Architects maintaining enterprise RAG search engines.
  • MLOps and Data Platform teams seeking reliable vector pipeline orchestration.
  • Compliance and Legal teams requiring strict right-to-erasure and policy deprecation guarantees in AI systems.

Why orchestrate this with Kestra

Maintaining bidirectional consistency between raw file storage (Git, S3, GCS) and vector databases requires multi-technology coordination: file scanning, cryptographic hashing, in-memory relational SQL adjudication, vector API operations, and notification hooks. Kestra executes this entire pipeline with zero external server dependencies, full observability, and native audit trail storage.

Prerequisites

  • Access to your source document repository or object storage bucket.
  • Target vector database (Pinecone, Qdrant, Weaviate, or pgvector).
  • Slack incoming webhook URL for reconciliation notifications.

Secrets

  • SLACK_WEBHOOK_URL: Slack incoming webhook endpoint for operational notifications.
  • DOCS_WEBHOOK_KEY: Secret key securing event-driven documentation release webhooks.

Quick start

  1. Configure SLACK_WEBHOOK_URL and DOCS_WEBHOOK_KEY in your Kestra namespace secrets.
  2. Import this blueprint into your Kestra instance.
  3. Click Execute to run the reconciliation flow against the included reference knowledge base.
  4. Observe the DuckDB reconciliation output, generated audit trail artifact, and Slack alert.
  5. Connect the webhook trigger to your Git repository's documentation release workflow or enable the nightly schedule.

How to extend

  • Swap the Python scanning script to ingest directly from io.kestra.plugin.aws.s3.List or io.kestra.plugin.gcp.gcs.List.
  • Connect io.kestra.plugin.pinecone.DeleteVectors and io.kestra.plugin.pinecone.UpsertVectors directly in downstream tasks.
  • Add semantic similarity benchmarking after re-embedding to ensure retrieval accuracy.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.