Schedule icon
Script icon
Process icon
AwsCLI icon
If icon
SlackIncomingWebhook icon
Log icon

AWS Cost Explorer Spend Anomaly Guard

Watch AWS spend while the month runs. Query Cost Explorer daily, compare month-to-date vs the same window last month, and alert Slack when costs spike.

Categories
CloudInfrastructure

Cloud bills surprise people because spend is invisible between invoices: a runaway batch job, a leaked NAT Gateway, or a traffic spike can run for weeks before anyone sees a number. This blueprint checks the pace of spending while the month is still running. It queries the AWS Cost Explorer API for month-to-date UnblendedCost grouped by service, queries the identical day-range one month earlier for a like-for-like baseline, computes the percent change in Python, and lets an If task compare it against your threshold — over it, Slack gets the measured numbers and the top service with time left to react.

How it works

  1. compute_windows (io.kestra.plugin.scripts.python.Script on the io.kestra.plugin.core.runner.Process runner) builds two windows: this month from the 1st through tomorrow (Cost Explorer's End is exclusive), and the same number of days starting at the 1st of the previous month. Comparing identical day-ranges avoids the classic mistake of judging a partial month against a full one.
  2. fetch_current_by_service (io.kestra.plugin.aws.cli.AwsCLI) runs aws ce get-cost-and-usage for the current window with --granularity DAILY --metrics UnblendedCost --group-by Type=DIMENSION,Key=SERVICE, emitting a flat [service, amount] list through stdOut.
  3. fetch_previous_by_service runs the identical query for the earlier window, producing the baseline in the same shape.
  4. compare_spend (Python) totals both payloads per service, rounds the month-to-date and baseline totals to cents, computes the percent change against a $0.01 floor — so brand-new spend against an empty baseline always trips the guard — and picks the leading service from the current window.
  5. check_spike (io.kestra.plugin.core.flow.If) compares pct_change with the spike_threshold_pct input. Over it, alert_spike posts both totals, the percent change, and the top service to Slack; under it, log_on_track keeps a quiet record so the execution history shows the month's trajectory.
  6. The errors block alerts Slack when the guard itself fails, and a daily Schedule trigger (shipped disabled) drives the cadence once enabled.

What you get

  • A like-for-like spend trajectory: month-to-date versus the same days last month, not a partial-versus-full comparison.
  • Per-service attribution in the alert, so the message names where to look instead of just that something changed.
  • A threshold that lives in an input, adjustable per environment without editing the flow.
  • A cost_summary JSON output (current cost, baseline, percent change, top service) for dashboards or downstream flows.

Who it's for

  • FinOps engineers who own the AWS bill and need a mid-month alarm on spending pace.
  • Platform and SRE teams watching for runaway workloads before the invoice arrives.
  • Engineering leaders who have been surprised by an AWS invoice at least once.

Why orchestrate this with Kestra

Cost Explorer exposes the numbers, but an API nobody polls prevents nothing. Kestra provides the loop around it: a schedule, retries on flaky API calls, a Python step for the window arithmetic with no infrastructure of its own thanks to the Process runner, a branch that separates alerting from quiet record-keeping, and an execution history that becomes a day-by-day audit of the month. Plain cron scripts give you none of the observability, replay, or alert routing that orchestration adds around the same two CLI calls.

Prerequisites

  • An AWS account with Cost Explorer enabled and an IAM principal with ce:GetCostAndUsage. Cost Explorer data can lag up to 24 hours, so very early-month runs compare against partially populated data.
  • The ce endpoint is only served from us-east-1 and us-west-2 — keep the aws_region input on one of those.
  • A Slack incoming webhook for cost alerts.
  • Python 3 available on the Kestra worker, since both script steps use the Process runner; switch them to a Docker task runner if the host has no Python.

Secrets

  • AWS_ACCESS_KEY_ID: AWS access key with ce:GetCostAndUsage permission.
  • AWS_SECRET_ACCESS_KEY: secret key matching the access key.
  • SLACK_WEBHOOK_URL: Slack incoming webhook URL for spike and failure alerts.

Quick start

  1. Add the three secrets to your Kestra namespace.
  2. Execute the flow once and read cost_summary (and the logs) to see your real month-to-date trajectory.
  3. Set spike_threshold_pct to a sensible margin above your normal month-over-month variance.
  4. Set disabled: false on the daily_cost_check trigger.

How to extend

  • Compare against the previous full month by changing compute_windows, or against a 30-day rolling window for always-on services.
  • Add a second If on an absolute floor (for example, alert above $50 even when the percentage is calm) to catch slow, steady creep.
  • Break out cost by account or region with --group-by Type=DIMENSION,Key=LINKED_ACCOUNT instead of SERVICE.
  • Fan the alert out to PagerDuty or a FinOps channel by adding another notification task in the then branch.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.