New to Kestra?
Use blueprints to kickstart your first workflows.
Build and incrementally update a Neo4j knowledge graph from GitHub Markdown docs with Claude and Kestra, then answer questions with cited sources.
Vector search finds passages that sound similar to a question. It struggles with questions about how things connect, such as which tasks depend on a feature, or which alternatives exist for a concept. GraphRAG answers those by storing entities and their relationships in a graph and walking it. The hard part is not the query. It's keeping the graph built and current as the documentation changes. This blueprint automates that pipeline. It reads a folder of Markdown docs from GitHub, uses Claude to extract entities and relationships, merges them into Neo4j with safe, parameterized Cypher, re-processes only files that changed, and answers a question from the graph with the source documents cited.
ensure_constraint (io.kestra.plugin.neo4j.Query) creates a uniqueness constraint on entity keys, so repeated or overlapping runs never create duplicate nodes.list_files (io.kestra.plugin.core.http.Request) lists every file under docs_path with one call to GitHub's git trees API. Each file comes with its Git blob SHA, so no repository clone is needed, even for very large repositories.known_docs (io.kestra.plugin.neo4j.Query) reads the SHA stored on each Document node. The graph itself is the record of what was ingested, so there is no separate state to keep in sync.select_docs (io.kestra.plugin.scripts.python.Script) keeps the .md and .mdx files that are new or whose SHA changed, capped at max_files per run. ingest_changed (io.kestra.plugin.core.flow.If) skips extraction entirely when nothing changed, so an unchanged repository costs no model calls.ingest_docs (io.kestra.plugin.core.flow.Loop) processes each selected file, one at a time:download_doc (io.kestra.plugin.core.http.Download) fetches the raw Markdown. Raw downloads don't count against the GitHub API rate limit.extract (io.kestra.plugin.ai.completion.ChatCompletion with Claude) returns entities with a type and description, and relationships with a type and a short evidence quote, as JSON.entity_rows and relationship_rows (io.kestra.plugin.core.storage.Write) turn the reply into JSON arrays, tolerating code fences, and io.kestra.plugin.serdes.json.JsonToIon converts them to rows.clear_previous_facts removes the edges and mentions this document produced last time, so a changed file replaces its old facts instead of piling up.load_entities and load_relationships (io.kestra.plugin.neo4j.Batch) merge the rows with UNWIND $props. Extracted text always arrives as query parameters and never becomes Cypher. Entities merge on a lower-cased key, and each relationship is stored once per source document, so every fact keeps its own citation.mark_ingested records the file's SHA only after both loads succeed, so a failed file is retried on the next run.graph_stats counts the documents, entities and relationships in the graph.question_terms asks Claude for the key terms in question. retrieve_subgraph slugifies each term, so only letters, digits and hyphens reach the Cypher text, then returns the matching entities' relationships with evidence and source documents.answer (Claude) answers the question using only those graph facts and cites the source documents. If the graph doesn't hold the answer, it says what's missing instead of guessing.on_docs_push trigger (io.kestra.plugin.core.trigger.Webhook) lets a GitHub push webhook start an incremental update whenever the docs change.$props parameters, and question terms are slugified before they touch Cypher.answer flow output, plus graph_size and ingestion outputs for monitoring.max_files, so large documentation sets are ingested over several runs.Extracting entities with an LLM is one call. Running it as a pipeline is what's hard: deciding what changed, keeping each document's facts replaceable, loading safely, retrying a failed file without reprocessing everything, and capping cost. Kestra handles the change detection and per-file Loop, keeps secrets out of the flow, runs Neo4j loads as parameterized batches, records each run's inputs and outputs, and lets a GitHub webhook or a schedule keep the graph current. If one file fails, the next run picks it up automatically because its SHA was never recorded.
docker run -p 7474:7474 -p 7687:7687 -e NEO4J_AUTH=neo4j/<password> neo4j:5, or Neo4j AuraDB.anthropic-workspace-id header that keys without a workspace require.NEO4J_URL: Bolt URL of the database, for example bolt://neo4j:7687 or neo4j+s://<id>.databases.neo4j.io for AuraDB.NEO4J_USERNAME: Neo4j user.NEO4J_PASSWORD: Neo4j password.ANTHROPIC_API_KEY: Anthropic API key used for extraction, question terms and the answer.GRAPHRAG_WEBHOOK_KEY: secret key in the webhook URL that GitHub calls.repo (STRING, default kestra-io/docs): public repository as owner/name.branch (STRING, default main): branch to read.docs_path (STRING, default src/contents/docs/05.workflow-components): folder to ingest. Every .md and .mdx file below it is included.max_files (INT, default 10): new or changed files sent to the model per run. Set it to 0 to only answer a question from the existing graph.question (STRING, default Which tasks can a flow use to wait for a condition or for a person before it continues?): answered from the graph at the end of the run.{{ outputs.answer }} (flow output) / {{ outputs.answer.textOutput }}: the cited answer.{{ outputs.graph_size }} (flow output) / {{ outputs.graph_stats.row }}: documents, entities and relationships in the graph.{{ outputs.ingestion }} (flow output): total, changed and processed files for the run.{{ outputs.select_docs.vars.paths }}: the files processed in this run.{{ outputs.retrieve_subgraph.rows }}: the graph facts used for the answer, each with evidence and source.MATCH (e:Entity)-[r:RELATES_TO]->(n) RETURN e, r, n LIMIT 100 to explore the graph./api/v1/main/executions/webhook/company.team/graphrag-docs-to-neo4j-knowledge-graph/<GRAPHRAG_WEBHOOK_KEY>.repo and docs_path at your own docs, runbooks or architecture decision records.io.kestra.plugin.scrapy.CLI, PDFs with io.kestra.plugin.tika.Parse, or use Apify for managed scraping, then feed the text to extract.claude_model to claude-sonnet-5-5 or claude-haiku-4-5 for extraction, or swap in another provider by changing provider.concurrencyLimit once the uniqueness constraint is in place, and add a retry on the load tasks.io.kestra.plugin.core.trigger.Schedule as a fallback to the webhook, or a cleanup step that removes Document nodes for files deleted from the repository.MERGE design.Neo4j plugin: https://kestra.io/plugins/plugin-neo4j
AI plugin: https://kestra.io/plugins/plugin-ai
Loop task: https://kestra.io/plugins/core/tasks/flow/io.kestra.plugin.core.flow.loop
GitHub git trees API: https://docs.github.com/en/rest/git/trees
Neo4j Cypher MERGE: https://neo4j.com/docs/cypher-manual/current/clauses/merge/
This blueprint has been created by Biplab Bera
