Schedule icon
Webhook icon
Script icon
Process icon
If icon
Create icon
SlackIncomingWebhook icon
Log icon

Catch Link Rot Before Your Users and Your Crawler Do

Crawl your sitemap, check every indexed URL, and get a GitHub issue and Slack alert listing the pages that return errors or time out.

Categories
BusinessData

Every site accumulates link rot: pages move, deploys drop routes, redirects expire, and the sitemap keeps listing what used to exist. The audit that would catch it is exactly the chore nobody schedules — until an SEO report or a customer does it for you. This blueprint reads your sitemap (index files and gzipped sitemaps included), probes each URL with a HEAD request and a GET fallback so no CDN quirks create false alarms, and when anything fails it opens a GitHub issue with the full URL/status table and tells Slack. Clean runs log quietly, so the execution history doubles as proof the audit keeps happening.

How it works

  1. crawl_sitemap (io.kestra.plugin.scripts.python.Script on the io.kestra.plugin.core.runner.Process runner) fetches the sitemap, decompresses .gz, follows up to five child sitemaps when it is an index, and caps the crawl at max_urls. Every URL is probed concurrently (8 workers) with HEAD, falling back to GET for servers that reject HEAD, treating HTTP errors and timeouts as broken. checked, broken_count, and the first 25 broken entries are emitted through the ::{"outputs": ...}:: protocol. A sitemap with zero URLs fails fast rather than reporting a clean audit.
  2. check_links (io.kestra.plugin.core.flow.If) branches on broken_count. Over zero: create_issue (io.kestra.plugin.github.issues.Create) opens a labeled issue with a markdown table of every failing URL, and alert_slack posts the counts; at zero: log_clean records the all-clear.
  3. The errors block alerts Slack when the audit itself breaks — an unreachable sitemap must never read as zero broken links.
  4. Triggers: a weekly Schedule (shipped disabled) plus a Webhook (sitemap-audit) to re-check right after a deploy or migration.

What you get

  • An automated link-rot audit: every indexed URL checked on a cadence, no browser extensions, no cron glue.
  • False-alarm-resistant probing: HEAD first, GET fallback, timeouts counted as broken.
  • A GitHub issue as the durable record — the failing pages land in the team's existing backlog with labels.
  • An audit_summary JSON output (checked, broken_count, broken) for dashboards or a release gate.

Who it's for

  • Content and docs teams whose sites outgrow manual link checks.
  • SEO-minded engineers who want sitemap hygiene covered by the same orchestrator that runs everything else.
  • Agencies and publishers managing sites where dead pages cost money.

Why orchestrate this with Kestra

A link checker script finds broken links; it does not decide weekly cadence vs post-deploy trigger, open the backlog item, notify the channel, alert on its own failure, and keep a history proving the audit ran. Wrapping the crawl in a flow makes the audit event-driven (the webhook), durable (the issue), and self-reporting (the errors alert) — and the audit_summary output makes "no broken links" a condition other flows can require.

Prerequisites

  • A public (or Worker-reachable) sitemap URL.
  • A GitHub token with issue-create permission on the repository input.
  • A Slack incoming webhook for failure and broken-link alerts.

Secrets

  • GITHUB_TOKEN: GitHub token used to open the broken-links issue.
  • SLACK_WEBHOOK_URL: Slack incoming webhook URL for audit and failure alerts.

Quick start

  1. Add both secrets to your Kestra namespace.
  2. Set sitemap_url to your sitemap and repository to the target repo.
  3. Run once and read audit_summary — expect a sensible checked count and your known-bad pages (if any) listed.
  4. Enable the weekly_audit schedule, and call the sitemap-audit webhook after deploys.

How to extend

  • Require zero broken links before a release by chaining this flow as a subflow in your deploy pipeline and branching on audit_summary.broken_count.
  • Recreate the GitHub issue instead of opening a new one each week by searching open broken-links issues first with io.kestra.plugin.github.issues.Search.
  • Extend the probe set: fetch each broken URL's redirects and report the redirect chain that died.
  • Push the audit_summary into a dashboard alongside uptime and performance checks.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.