Schedule icon
Loop icon
Download icon
Parse icon
Return icon
If icon
SlackIncomingWebhook icon
Log icon

Document Ingestion Quality Gate with Apache Tika and Slack Alerts

Automatically extract text with Apache Tika, validate MIME type and character thresholds, and send Slack alerts for failing documents.

Categories
CoreData

Automate pre-ingestion document quality validation before downstream vector embedding or LLM pipelines. This blueprint streams documents from web URLs into internal storage, extracts text and metadata via embedded Apache Tika, normalizes Content-Type headers, and applies quality gates for character length and MIME type. When a document fails validation, an alert is posted to Slack.

How it works

  1. weekly_schedule (io.kestra.plugin.core.trigger.Schedule) triggers weekly execution (disabled: true).
  2. process_documents (io.kestra.plugin.core.flow.Loop) loops over target document URLs.
  3. download_document (io.kestra.plugin.core.http.Download) streams each document into internal storage.
  4. parse_document (io.kestra.plugin.tika.Parse) extracts text and metadata inline with contentType: TEXT.
  5. normalize_mime (io.kestra.plugin.core.debug.Return) strips parameters from Tika's Content-Type metadata header.
  6. check_quality (io.kestra.plugin.core.flow.If) evaluates character length against inputs.min_chars and checks MIME against allowed types (application/pdf, text/plain, text/html).
  7. Failing documents trigger alert_quality_failure (SlackIncomingWebhook), while passing files emit log_quality_pass.

What you get

  • Pre-ingestion validation preventing empty or corrupt files from entering vector stores.
  • Embedded Tika parsing for PDFs, HTML, and plain text with zero extra containers.
  • Automated Slack incident notifications detailing URL, character count, and MIME type.

Who it's for

  • Data engineers building RAG or LLM document ingestion pipelines.
  • Content operations teams auditing document publishing pipelines.

Why orchestrate this with Kestra

Manual document scripts lack state management and retry resilience. Kestra orchestrates fetch, parse, normalization, and conditional branching declaratively with full execution logging and secret management.

Prerequisites

  • Outbound HTTP access to document source URLs.
  • Slack Incoming Webhook URL configured as a secret.

Secrets

  • SLACK_WEBHOOK_URL: Slack Incoming Webhook URL. Open-source deployments configure SECRET_SLACK_WEBHOOK_URL in environment variables.

Quick start

  1. Add SLACK_WEBHOOK_URL secret to your Kestra namespace.
  2. Run flow with default inputs to process sample documents.
  3. Inspect execution logs and verify Slack alert for dummy.pdf.

Inputs

  • document_urls (ARRAY of STRING): List of target document URLs. Default: the Hugging Face app_store.pdf, RFC 2324 text, Python license HTML, and the W3C dummy.pdf (deliberate failing case).
  • min_chars (SELECT): Minimum character threshold (500, 1000, 5000; default "1000").

Outputs

  • download_document.uri: Internal storage URI at {{ outputs.download_document.uri }}.
  • parse_document.result.content: Extracted string inline text at {{ outputs.parse_document.result.content }}.
  • normalize_mime.value: Cleaned MIME string at {{ outputs.normalize_mime.value }}.

Pitfalls

  • A network failure or HTTP 404 on download_document aborts the flow and invokes the errors block.
  • MIME validation uses Tika parser metadata rather than raw HTTP server response headers.
  • Scanned PDFs without an OCR text layer fail the character threshold (OCR is disabled).

How to extend

  • Route passing documents to vector database embedding tasks or S3 storage.
  • Add custom MIME types to the allowed list in check_quality.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.