Schedule icon
Webhook icon
Script icon
Process icon
If icon
SlackIncomingWebhook icon
Log icon

Audit LLM Endpoints for Prompt Injections and Canary Leaks

Stress-test your LLM against prompt injection, DAN jailbreaks, and canary token leakage, generating an audit report and alerting Slack on breaches.

Categories
AICore

As generative AI systems are integrated into production workflows, prompt injections and jailbreak attacks present severe security risks. Attackers can manipulate model instructions to extract proprietary system prompts, bypass corporate safety filters, or leak sensitive API tokens.

This blueprint provides an automated, repeatable security evaluation harness for any OpenAI-compatible LLM endpoint. It simulates adversarial attacks—including direct instruction overrides, persona jailbreaks (DAN), delimiter confusion, and encoding obfuscation—against a system prompt containing a secret canary token. It automatically compiles a downloadable Markdown security audit report and triggers an immediate Slack security alert with the exact compromised vector IDs if a breach is detected.

How it works

  1. run_guardrail_audit (io.kestra.plugin.scripts.python.Script on the io.kestra.plugin.core.runner.Process runner) iterates through a configurable suite of adversarial prompts against your model endpoint. Each query includes an enterprise protection system prompt with an embedded verification canary. Responses are inspected for canary token leakage or jailbreak confirmation. Metrics and per-case test results are emitted via Kestra's stdout outputs protocol.
  2. generate_audit_report (io.kestra.plugin.scripts.python.Script) consumes the audit results and synthesizes a structured Markdown compliance report file (security-guardrail-report.md), persisted into Kestra internal storage with an executive summary and full attack vector breakdown table.
  3. evaluate_security_gate (io.kestra.plugin.core.flow.If) evaluates whether bypasses_detected exceeds max_allowed_bypasses. If compromised, alert_security_breach posts an alert with failing attack IDs and report URI to Slack; otherwise, log_guardrail_passed logs the clean audit.
  4. The errors block alerts Slack if the audit harness fails to connect or authenticate, ensuring broken credentials are never mistaken for a secure system.
  5. Triggers: a weekly Schedule (shipped disabled) plus an event-based Webhook to run automated security regression checks in CI/CD pipelines whenever system prompts or model versions change.

What you get

  • Automated security testing against known prompt injection and jailbreak techniques.
  • Zero external Python package dependencies (runs on Python standard library).
  • Canary token leak detection to identify unauthorized prompt disclosures.
  • An automatically generated, downloadable Markdown security audit report (security-guardrail-report.md).
  • An audit_summary JSON output with total tests, bypass counts, breach rate, and compromised vector IDs.
  • Immediate Slack alerting with specific attack vector IDs for rapid remediation.

Who it's for

  • AI Engineers deploying LLM applications who need to verify model robustness before shipping.
  • Security teams and DevSecOps engineers auditing model endpoints for compliance and data leakage.
  • Platform engineers managing local models (Ollama, vLLM) or proprietary cloud LLMs.

Why orchestrate this with Kestra

Securing LLM applications requires continuous testing rather than one-off checks. Models update, temperature and system instructions change, and new prompt variations emerge. Kestra provides scheduled audits, webhook triggers for CI/CD gates, failure handling, downloadable artifact reporting, and Slack notifications, while recording audit score trends across all execution runs.

Prerequisites

  • An OpenAI-compatible chat completions endpoint (OpenAI, Ollama, vLLM, or LocalAI).
  • A Slack incoming webhook for security alerts.

Secrets

  • OPENAI_API_KEY: API key for model authentication (can be a dummy string for local Ollama endpoints).
  • SLACK_WEBHOOK_URL: Slack webhook URL for security breach notifications.

Quick start

  1. Configure the OPENAI_API_KEY and SLACK_WEBHOOK_URL secrets in your Kestra namespace.
  2. Run the flow manually to establish a security baseline.
  3. Review the audit_summary output and download the security-guardrail-report.md artifact from the Kestra execution page.
  4. Enable the weekly_security_audit schedule or trigger via the audit_webhook in your CI/CD pipeline.

How to extend

  • Add custom attack vectors to the attack_suite input to test domain-specific sensitive data patterns.
  • Integrate with an automated ticket creation task (e.g. Jira or GitHub Issues) when critical breaches occur.
  • Add a secondary LLM judge task to evaluate semantic refusal quality.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.