New to Kestra?
Use blueprints to kickstart your first workflows.
Automatically extract text with Apache Tika, validate MIME type and character thresholds, and send Slack alerts for failing documents.
Automate pre-ingestion document quality validation before downstream vector embedding or LLM pipelines. This blueprint streams documents from web URLs into internal storage, extracts text and metadata via embedded Apache Tika, normalizes Content-Type headers, and applies quality gates for character length and MIME type. When a document fails validation, an alert is posted to Slack.
weekly_schedule (io.kestra.plugin.core.trigger.Schedule) triggers weekly execution (disabled: true).process_documents (io.kestra.plugin.core.flow.Loop) loops over target document URLs.download_document (io.kestra.plugin.core.http.Download) streams each document into internal storage.parse_document (io.kestra.plugin.tika.Parse) extracts text and metadata inline with contentType: TEXT.normalize_mime (io.kestra.plugin.core.debug.Return) strips parameters from Tika's Content-Type metadata header.check_quality (io.kestra.plugin.core.flow.If) evaluates character length against inputs.min_chars and checks MIME against allowed types (application/pdf, text/plain, text/html).alert_quality_failure (SlackIncomingWebhook), while passing files emit log_quality_pass.Manual document scripts lack state management and retry resilience. Kestra orchestrates fetch, parse, normalization, and conditional branching declaratively with full execution logging and secret management.
SLACK_WEBHOOK_URL: Slack Incoming Webhook URL. Open-source deployments configure SECRET_SLACK_WEBHOOK_URL in environment variables.SLACK_WEBHOOK_URL secret to your Kestra namespace.dummy.pdf.document_urls (ARRAY of STRING): List of target document URLs. Default: the Hugging Face app_store.pdf, RFC 2324 text, Python license HTML, and the W3C dummy.pdf (deliberate failing case).min_chars (SELECT): Minimum character threshold (500, 1000, 5000; default "1000").download_document.uri: Internal storage URI at {{ outputs.download_document.uri }}.parse_document.result.content: Extracted string inline text at {{ outputs.parse_document.result.content }}.normalize_mime.value: Cleaned MIME string at {{ outputs.normalize_mime.value }}.download_document aborts the flow and invokes the errors block.check_quality.