Webhook icon
Query icon
Request icon
Script icon
If icon
Loop icon
Download icon
ChatCompletion icon
Anthropic icon
Write icon
JsonToIon icon
Batch icon

GraphRAG Knowledge Graph from GitHub Docs with Neo4j and Claude

Build and incrementally update a Neo4j knowledge graph from GitHub Markdown docs with Claude and Kestra, then answer questions with cited sources.

Categories
AIData

Vector search finds passages that sound similar to a question. It struggles with questions about how things connect, such as which tasks depend on a feature, or which alternatives exist for a concept. GraphRAG answers those by storing entities and their relationships in a graph and walking it. The hard part is not the query. It's keeping the graph built and current as the documentation changes. This blueprint automates that pipeline. It reads a folder of Markdown docs from GitHub, uses Claude to extract entities and relationships, merges them into Neo4j with safe, parameterized Cypher, re-processes only files that changed, and answers a question from the graph with the source documents cited.

How it works

  1. ensure_constraint (io.kestra.plugin.neo4j.Query) creates a uniqueness constraint on entity keys, so repeated or overlapping runs never create duplicate nodes.
  2. list_files (io.kestra.plugin.core.http.Request) lists every file under docs_path with one call to GitHub's git trees API. Each file comes with its Git blob SHA, so no repository clone is needed, even for very large repositories.
  3. known_docs (io.kestra.plugin.neo4j.Query) reads the SHA stored on each Document node. The graph itself is the record of what was ingested, so there is no separate state to keep in sync.
  4. select_docs (io.kestra.plugin.scripts.python.Script) keeps the .md and .mdx files that are new or whose SHA changed, capped at max_files per run. ingest_changed (io.kestra.plugin.core.flow.If) skips extraction entirely when nothing changed, so an unchanged repository costs no model calls.
  5. ingest_docs (io.kestra.plugin.core.flow.Loop) processes each selected file, one at a time:
    • download_doc (io.kestra.plugin.core.http.Download) fetches the raw Markdown. Raw downloads don't count against the GitHub API rate limit.
    • extract (io.kestra.plugin.ai.completion.ChatCompletion with Claude) returns entities with a type and description, and relationships with a type and a short evidence quote, as JSON.
    • entity_rows and relationship_rows (io.kestra.plugin.core.storage.Write) turn the reply into JSON arrays, tolerating code fences, and io.kestra.plugin.serdes.json.JsonToIon converts them to rows.
    • clear_previous_facts removes the edges and mentions this document produced last time, so a changed file replaces its old facts instead of piling up.
    • load_entities and load_relationships (io.kestra.plugin.neo4j.Batch) merge the rows with UNWIND $props. Extracted text always arrives as query parameters and never becomes Cypher. Entities merge on a lower-cased key, and each relationship is stored once per source document, so every fact keeps its own citation.
    • mark_ingested records the file's SHA only after both loads succeed, so a failed file is retried on the next run.
  6. graph_stats counts the documents, entities and relationships in the graph.
  7. question_terms asks Claude for the key terms in question. retrieve_subgraph slugifies each term, so only letters, digits and hyphens reach the Cypher text, then returns the matching entities' relationships with evidence and source documents.
  8. answer (Claude) answers the question using only those graph facts and cites the source documents. If the graph doesn't hold the answer, it says what's missing instead of guessing.
  9. The on_docs_push trigger (io.kestra.plugin.core.trigger.Webhook) lets a GitHub push webhook start an incremental update whenever the docs change.

What you get

  • A Neo4j knowledge graph of the entities in your docs and how they relate, with every relationship carrying an evidence quote and its source document.
  • Incremental updates driven by Git SHAs. Unchanged files are never sent to the model again, and changed files replace their previous facts.
  • Injection-safe loading. Model output reaches Neo4j only as $props parameters, and question terms are slugified before they touch Cypher.
  • A grounded answer to a question, with citations, as the answer flow output, plus graph_size and ingestion outputs for monitoring.
  • Bounded cost per run through max_files, so large documentation sets are ingested over several runs.

Who it's for

  • Teams building GraphRAG or knowledge-graph search over product docs, runbooks or internal wikis kept in Git.
  • Developer relations and documentation teams who want to see how concepts in their docs connect.
  • Data and AI engineers who need a repeatable, observable pipeline to keep a graph in sync with a source.

Why orchestrate this with Kestra

Extracting entities with an LLM is one call. Running it as a pipeline is what's hard: deciding what changed, keeping each document's facts replaceable, loading safely, retrying a failed file without reprocessing everything, and capping cost. Kestra handles the change detection and per-file Loop, keeps secrets out of the flow, runs Neo4j loads as parameterized batches, records each run's inputs and outputs, and lets a GitHub webhook or a schedule keep the graph current. If one file fails, the next run picks it up automatically because its SHA was never recorded.

Prerequisites

  • A Neo4j 5 database reachable from Kestra, for example docker run -p 7474:7474 -p 7687:7687 -e NEO4J_AUTH=neo4j/<password> neo4j:5, or Neo4j AuraDB.
  • An Anthropic API key, created inside a workspace in the Claude Console. Kestra can't send the anthropic-workspace-id header that keys without a workspace require.
  • A public GitHub repository with Markdown docs. No GitHub token is needed. The flow makes one GitHub API call per run, well inside the unauthenticated limit of 60 per hour.
  • A Kestra worker with Docker available, for the Python selection step.

Secrets

  • NEO4J_URL: Bolt URL of the database, for example bolt://neo4j:7687 or neo4j+s://<id>.databases.neo4j.io for AuraDB.
  • NEO4J_USERNAME: Neo4j user.
  • NEO4J_PASSWORD: Neo4j password.
  • ANTHROPIC_API_KEY: Anthropic API key used for extraction, question terms and the answer.
  • GRAPHRAG_WEBHOOK_KEY: secret key in the webhook URL that GitHub calls.

Inputs

  • repo (STRING, default kestra-io/docs): public repository as owner/name.
  • branch (STRING, default main): branch to read.
  • docs_path (STRING, default src/contents/docs/05.workflow-components): folder to ingest. Every .md and .mdx file below it is included.
  • max_files (INT, default 10): new or changed files sent to the model per run. Set it to 0 to only answer a question from the existing graph.
  • question (STRING, default Which tasks can a flow use to wait for a condition or for a person before it continues?): answered from the graph at the end of the run.

Outputs

  • {{ outputs.answer }} (flow output) / {{ outputs.answer.textOutput }}: the cited answer.
  • {{ outputs.graph_size }} (flow output) / {{ outputs.graph_stats.row }}: documents, entities and relationships in the graph.
  • {{ outputs.ingestion }} (flow output): total, changed and processed files for the run.
  • {{ outputs.select_docs.vars.paths }}: the files processed in this run.
  • {{ outputs.retrieve_subgraph.rows }}: the graph facts used for the answer, each with evidence and source.

Quick start

  1. Start Neo4j and create the secrets above.
  2. Save the flow and run it with the defaults. The first run ingests 10 files from the Kestra workflow components docs and answers the default question.
  3. Run it again to ingest the next batch. Files already in the graph are skipped.
  4. Open the Neo4j Browser at http://localhost:7474 and run MATCH (e:Entity)-[r:RELATES_TO]->(n) RETURN e, r, n LIMIT 100 to explore the graph.
  5. To keep the graph current, add a GitHub webhook on the docs repository for push events, pointing at /api/v1/main/executions/webhook/company.team/graphrag-docs-to-neo4j-knowledge-graph/<GRAPHRAG_WEBHOOK_KEY>.

How to extend

  • Point repo and docs_path at your own docs, runbooks or architecture decision records.
  • Ingest websites with io.kestra.plugin.scrapy.CLI, PDFs with io.kestra.plugin.tika.Parse, or use Apify for managed scraping, then feed the text to extract.
  • Control cost by setting claude_model to claude-sonnet-5-5 or claude-haiku-4-5 for extraction, or swap in another provider by changing provider.
  • Run the extraction loop with a higher concurrencyLimit once the uniqueness constraint is in place, and add a retry on the load tasks.
  • Add a scheduled io.kestra.plugin.core.trigger.Schedule as a fallback to the webhook, or a cleanup step that removes Document nodes for files deleted from the repository.
  • Use FalkorDB or Memgraph instead of Neo4j through a script task, keeping the same JSON-to-MERGE design.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.