Schedule icon
Query icon
Return icon
Switch icon
Log icon
SlackIncomingWebhook icon

Cassandra Keyspace Replication Audit Gate

Audit Cassandra keyspace replication strategy via CQL on a schedule and alert Slack when SimpleStrategy is found or the cluster is unreachable.

Categories
DataInfrastructure

SimpleStrategy ignores datacenter topology and is Cassandra's own documented default for anything other than a single-node development cluster, yet a keyspace created without specifying NetworkTopologyStrategy silently keeps it in production. This blueprint reads system_schema.keyspaces on a schedule, flags any non-system keyspace still using SimpleStrategy, and pages Slack the moment one is found or the cluster cannot be queried at all. plugin-cassandra has no dedicated cluster-health or topology task, so its native CQL query task against Cassandra's own schema metadata is the audit.

How it works

  1. query_keyspace_replication (io.kestra.plugin.cassandra.standard.Query) runs SELECT keyspace_name, replication FROM system_schema.keyspaces with fetchType: FETCH and allowFailure: true, so a connection or authentication failure still lets the flow continue instead of aborting.
  2. classify_replication_risk (io.kestra.plugin.core.debug.Return) checks whether outputs.query_keyspace_replication.rows is defined; if not, UNREACHABLE. Otherwise it runs a jq filter over the rows that counts keyspaces whose replication.class contains SimpleStrategy, skipping names starting with system when exclude_system_keyspaces is true, and returns AT_RISK when that count is above zero, HEALTHY otherwise.
  3. route_by_risk (io.kestra.plugin.core.flow.Switch) branches on that string: HEALTHY logs a quiet confirmation, AT_RISK pages Slack naming the host and the exclusion mode, UNREACHABLE pages Slack separately, and defaults logs the raw outcome.
  4. log_audit_status records a status line on every run regardless of branch, so the execution history is a continuous compliance trend.
  5. The errors block pages Slack separately if the flow itself fails outright.
  6. The cassandra_replication_audit_schedule trigger polls every 6 hours with explicit input overrides, shipped disabled: true until you point it at a real cluster.

What you get

  • A scheduled CQL-native replication-strategy audit with zero custom client code.
  • A three-way status output (HEALTHY / AT_RISK / UNREACHABLE) you can chart or gate other flows on.
  • A Slack page naming the host, port, and exclusion mode behind an unsafe replication finding.
  • A separate failure channel for the audit flow itself, plus a standing audit log line per run.

Who it's for

  • Database reliability and platform teams running self-managed Cassandra clusters who want a recurring check that new keyspaces were not created with the wrong replication strategy.
  • Teams migrating from a single-node dev cluster to a multi-node or multi-DC deployment, where leftover SimpleStrategy keyspaces are an easy miss.

Why orchestrate this with Kestra

A one-off DESCRIBE KEYSPACES from cqlsh proves the setup is correct for one moment. This flow keeps every audit in the execution history, turns the result into a typed output other flows can depend on, and gives you a Switch branch to grow into real remediation (opening a ticket, running an ALTER KEYSPACE) without rewriting the audit itself.

Prerequisites

  • A running Cassandra cluster with its native CQL port reachable from the Kestra Worker. For local testing: docker run -d --name cassandra -p 9042:9042 cassandra:latest, then connect it to your Kestra worker's Docker network so cassandra_host resolves: docker network connect YOUR_KESTRA_NETWORK cassandra (run docker network ls to find the real network name). Allow a minute or two for the single node to finish bootstrapping before the first run.
  • A Slack incoming webhook for the at-risk and unreachable alerts.
  • If your cluster requires authentication, query_keyspace_replication.session accepts username and password properties, not used here since the default Cassandra image runs unauthenticated.

Secrets

  • SLACK_WEBHOOK_URL: Slack incoming webhook used for the at-risk alert, the unreachable alert, and the flow-failure alert.
  • OSS note: if your Kestra instance reads secrets from environment variables instead of a secrets backend, set SECRET_SLACK_WEBHOOK_URL as base64, for example echo -n "<your-webhook-url>" | base64.

Inputs

  • cassandra_host (STRING, default cassandra): hostname or IP of a Cassandra contact point.
  • cassandra_port (INT, default 9042): native CQL port.
  • cassandra_datacenter (STRING, default datacenter1): local datacenter name for the driver's load-balancing policy.
  • exclude_system_keyspaces (SELECT, one of "true", "false", default "true"): skip Cassandra's own system keyspaces, which use SimpleStrategy by design.

Quick start

  1. Add the SLACK_WEBHOOK_URL secret.
  2. Set cassandra_host and cassandra_port to your contact point.
  3. Run the flow once on a fresh single-node cluster and confirm status reads HEALTHY (its only keyspaces are system ones, excluded by default).
  4. Create a test keyspace with SimpleStrategy and re-run to confirm status reads AT_RISK.
  5. Enable cassandra_replication_audit_schedule.

Outputs

  • status (STRING): the classified outcome, {{ outputs.classify_replication_risk.value }}, also exposed as the flow-level output {{ outputs.status }}.
  • {{ outputs.query_keyspace_replication.rows }}: the raw list of keyspace_name/replication maps, when the query succeeded.

Pitfalls

  • A fresh single-node cluster has no user keyspaces at all, only system ones; with exclude_system_keyspaces: true that audits zero rows and always reads HEALTHY. Create a real user keyspace to exercise the AT_RISK branch.
  • Setting exclude_system_keyspaces to false will almost always report AT_RISK on a default Cassandra install, since several system keyspaces ship with SimpleStrategy by design; use false only for a full inventory view, not as the alerting default.
  • This audit only reads the class field; it does not evaluate whether a NetworkTopologyStrategy keyspace's per-datacenter replication factor is itself adequate, since those factors are stored under arbitrary datacenter-name keys rather than a fixed field.
  • A 6-hour poll can miss a keyspace that is created with SimpleStrategy and altered again before the next run.
  • allowFailure: true on the query means a total connection failure and a successful query returning zero matching rows both need to be told apart by checking whether rows is defined at all, not by its length.

How to extend

  • Also alert when durable_writes is false on a keyspace that should not tolerate data loss on an unclean shutdown.
  • Store the current AT_RISK keyspace list in a KV bucket and alert only on new entries instead of every run.
  • Parse NetworkTopologyStrategy's per-datacenter replication factors and gate on a minimum value once the datacenter names are known ahead of time.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.