Webhook icon
Log icon
Script icon
Docker icon
If icon
SlackIncomingWebhook icon

Multi-Provider LLM Fallback Router with Circuit Breaker

Route LLM requests to a primary model and automatically fail over to a secondary provider on rate limits or outages, ensuring high availability.

Categories
AICore

Production applications, customer-facing chatbots, and automated agent pipelines are vulnerable to upstream LLM disruptions:

  • HTTP 429 Rate Limits: Spikes in traffic quickly exhaust Token Per Minute (TPM) or Request Per Minute (RPM) tiers on OpenAI or Anthropic.
  • Latency Spikes & 503 Outages: Upstream provider degradation (p99 latency > 10s or Cloudflare gateway timeouts) causes downstream workflows to hang and fail.
  • Single Point of Failure: Hardcoding a single AI provider risks customer-facing downtime during provider outages.

This blueprint provides an enterprise-ready, high-availability LLM gateway:

  • Exposes an event-driven Webhook endpoint ready to receive inference requests from any application.
  • Attempts inference with the primary model (OpenAI gpt-4o-mini).
  • Implements an automated circuit breaker: if OpenAI returns HTTP 429, throws an internal error, or exceeds inputs.timeout_seconds, it automatically fails over to Google Gemini (gemini-1.5-flash) without failing the caller's request.
  • Emits normalized outputs (response_text, active_provider, fallback_triggered, latency_ms) so downstream applications receive consistent JSON schemas.
  • Dispatches an incident alert to Slack when failover occurs to inform AI platform engineers.

How it works

  1. llm_router_webhook receives incoming requests with prompt, system_prompt, temperature, and timeout_seconds.
  2. log_request logs generation parameters and caller details.
  3. route_inference launches a Python task to attempt primary inference on OpenAI. If an error or timeout occurs, it catches the exception and immediately requests generation from Google Gemini.
  4. check_failover (io.kestra.plugin.core.flow.If) branches:
    • If failover occurred: notify_slack_failover posts an incident card to Slack detailing the primary failure reason and active fallback model.
    • If primary succeeded: log_primary_success logs generation telemetry.
  5. An errors block catches catastrophic failures where both providers are unreachable.

What you get

  • 99.99% AI Availability: Protects user-facing services and background agents from single-provider downtime.
  • Automated Rate Limit Resilience: Seamlessly handles transient 429 Too Many Requests spikes without failing jobs.
  • Observability & Alerting: SREs and AI engineers receive instant Slack visibility whenever failover triggers.
  • Normalized API Contract: Callers interact with a single endpoint and receive identical response schemas regardless of which provider answered.

Prerequisites

  • An OpenAI API key.
  • A Google Gemini API key.
  • An incoming Slack webhook URL for alerts.

Secrets

  • OPENAI_API_KEY: API key for primary OpenAI model.
  • GEMINI_API_KEY: API key for secondary Google Gemini fallback model.
  • SLACK_WEBHOOK_URL: Slack Incoming Webhook URL.

Quick start

  1. Add OPENAI_API_KEY, GEMINI_API_KEY, and SLACK_WEBHOOK_URL to your Kestra namespace secrets.
  2. Trigger a manual execution from the Kestra UI to test normal generation via the primary provider.
  3. Simulate a failover by supplying an invalid key or setting timeout_seconds: 1 to observe the automatic failover to Gemini and the resulting Slack alert.
  4. Call the llm_router_webhook endpoint from your external services or subflows:
    curl -X POST "https://kestra.example.com/api/v1/executions/webhook/company.ai/ai-multi-provider-fallback-router/llm-router-endpoint" \
      -H "Content-Type: application/json" \
      -d '{
        "prompt": "Summarize today customer complaints into 3 action items.",
        "system_prompt": "You are a customer operations manager."
      }'
    

How to extend

  • Three-Tier Fallback: Add Anthropic Claude (claude-3-5-sonnet) as a third-tier fallback before failing.
  • Response Semantic Caching: Wrap the router with a Redis cache lookup (io.kestra.plugin.redis) to return cached answers for identical prompt hashes.
  • Cost & Token Logging: Record per-request token usage and estimated cost into a database table or BigQuery for FinOps tracking.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.