Schedule icon
Query icon
Return icon
Switch icon
Log icon
If icon
Loop icon
SlackIncomingWebhook icon

Cassandra Superuser Login Exposure Sentinel

Audit Cassandra for superuser roles that can log in directly, alert on exposure, and optionally revoke login with ALTER ROLE.

Categories
DataInfrastructure

A Cassandra role with is_superuser = true and no way to log in is inert - nobody can authenticate as it directly. The same role with can_login = true is a standing, directly usable, full-privilege credential, and nothing in the cluster itself tells you how many of those exist at any given moment. This blueprint queries system_auth.roles, flags any role that is both a superuser and able to log in, alerts on it, and can revoke that role's login capability with ALTER ROLE ... WITH LOGIN = false - leaving its SUPERUSER status, password, and every other permission untouched, so the role still exists for administrative purposes but cannot authenticate directly anymore.

How it works

  1. cassandra_superuser_audit_schedule (io.kestra.plugin.core.trigger.Schedule) runs every 6 hours, always forcing auto_remediate: "false" and dry_run: "true" through its own inputs: override regardless of this flow's defaults. Shipped disabled so you can validate a manual run first.
  2. query_superuser_roles (io.kestra.plugin.cassandra.standard.Query) runs SELECT role, is_superuser, can_login FROM system_auth.roles with fetchType: FETCH and allowFailure: true, so an unreachable cluster produces no rows instead of failing the task outright.
  3. classify_superuser_exposure (io.kestra.plugin.core.debug.Return) classifies the result as UNREACHABLE (no rows), AT_RISK (at least one role with both is_superuser and can_login true), or HEALTHY.
  4. route_by_exposure (io.kestra.plugin.core.flow.Switch) branches on that classification:
    • HEALTHY: log_healthy records a clean check.
    • AT_RISK: a nested route_remediation (Switch on auto_remediate) either considers remediation ("true", gated again by dry_run through check_dry_run) or only logs that remediation is disabled ("false"). When both gates allow it, revoke_superuser_login (io.kestra.plugin.core.flow.Loop) runs revoke_role_login (ALTER ROLE ... WITH LOGIN = false) for each flagged role. alert_at_risk always fires to Slack regardless of which path ran.
    • UNREACHABLE: alert_unreachable pages Slack directly - a broken audit must never read as "no exposure".
  5. log_audit_status always runs last, printing the classification regardless of which branch fired.
  6. The errors block alerts Slack separately if the flow itself fails outside these handled branches.

What you get

  • A direct answer to "how many logins-capable superusers exist right now", computed from Cassandra's own role table, not inferred from configuration files.
  • Remediation that only ever flips LOGIN to false; SUPERUSER status, the role's password, and its role memberships are never touched.
  • Two independent gates (auto_remediate, dry_run) between "exposure detected" and "login actually revoked".
  • A three-way classification (HEALTHY/AT_RISK/UNREACHABLE) and a final status log every run, so the execution history doubles as a privileged-access trend for the cluster.

Who it's for

  • Platform and security teams running self-hosted Cassandra who need a standing check on privileged credentials, not a one-time LIST ROLES review.
  • DBAs auditing a cluster after onboarding or offboarding administrators, to confirm no forgotten superuser login survived the change.
  • Anyone evaluating the Cassandra plugin who wants a worked example of system_auth.roles auditing alongside this repo's existing system_schema.keyspaces replication sentinel.

Why orchestrate this with Kestra

system_auth.roles is a live, read-only snapshot; querying it by hand with cqlsh answers one moment in time and fixes nothing. Kestra supplies the schedule, the three-way risk classification as a first-class branch, two independent authorization gates before anything is altered, and an execution history that shows exactly when a role became both a superuser and loginable, and whether it was revoked.

Prerequisites

  • A reachable Cassandra cluster and a role with rights to read system_auth.roles and, if remediation runs, execute ALTER ROLE - note that a role can alter another role's SUPERUSER status only if it is itself a superuser, and cannot alter the SUPERUSER/LOGIN status of the role it is currently authenticated as.
  • A Slack incoming webhook for alerts.

Local testing: docker run -d --name cassandra-superuser-sentinel -p 9043:9042 cassandra:latest docker network connect YOUR_KESTRA_NETWORK cassandra-superuser-sentinel (only needed if the Kestra Worker runs in a separate Docker network than this container; set cassandra_host to cassandra-superuser-sentinel from inside that network, or localhost with port 9043 from the host). The official image's default cassandra superuser role ships exactly in the audited state - is_superuser and can_login both true - so a first run against it is a guaranteed, realistic breach.

Secrets

  • SLACK_WEBHOOK_URL: Slack incoming webhook used by alert_at_risk, alert_unreachable, and the errors block.
  • This blueprint's session connects without a username/password, matching this repo's existing cassandra-replication-audit-gate.yaml pattern against an unauthenticated local instance. For an authenticated cluster, add the credential properties documented on io.kestra.plugin.cassandra.standard.Query's session block from {{ secret(...) }} values rather than hardcoding them.
  • In Kestra OSS (no Enterprise secrets backend), secrets are supplied as environment variables prefixed SECRET_, base64-encoded, and read back in flows with {{ secret('NAME') }} - for example SECRET_SLACK_WEBHOOK_URL=$(echo -n 'https://hooks.slack.com/...' | base64). This keeps credentials out of the flow YAML but, per Kestra's own documentation, offers no encryption at rest or access control beyond the host environment; use the Enterprise secrets backend for stronger guarantees.

Inputs

  • cassandra_host (STRING, default cassandra): Cassandra contact point hostname.
  • cassandra_port (INT, default 9042): native CQL port.
  • cassandra_datacenter (STRING, default datacenter1): local datacenter for the driver's load-balancing policy.
  • auto_remediate (SELECT: "false", "true"; default "false"): must be "true" for ALTER ROLE to even be considered.
  • dry_run (SELECT: "true", "false"; default "true"): must be explicitly "false", together with auto_remediate: "true", for login to actually be revoked.

Outputs

  • outputs.query_superuser_roles.rows: the raw role/is_superuser/can_login rows for every run (absent when the cluster is unreachable).
  • outputs.classify_superuser_exposure.value: HEALTHY, AT_RISK, or UNREACHABLE for every run.

Quick start

  1. Start Cassandra locally with the command above - the default cassandra role is already in the audited (exposed) state, no setup required.
  2. Add the SLACK_WEBHOOK_URL secret.
  3. Run the flow manually with the defaults and confirm alert_at_risk fires, and that log_remediation_disabled logs without altering anything.
  4. Run once with auto_remediate: "true" and dry_run: "true" and confirm log_dry_run_remediation describes the revocation without running it.
  5. Run once with auto_remediate: "true" and dry_run: "false" against a non-default role (see Pitfalls - the default cassandra role cannot revoke its own login while authenticated as itself), then confirm with SELECT role, can_login FROM system_auth.roles; that can_login is now false for that role.
  6. Enable cassandra_superuser_audit_schedule once you trust the check; it always runs in the safe auto_remediate: "false" / dry_run: "true" mode regardless of what you leave the flow's own defaults set to.

How to extend

  • Extend classify_superuser_exposure to also flag roles with a non-empty member_of that transitively inherits superuser status, which this audit's direct is_superuser check does not follow.
  • Loop over several cassandra_host values with io.kestra.plugin.core.flow.Loop (ForEach is deprecated) to sweep a multi-cluster fleet in one run.
  • Route the remediation branch through a Pause approval task instead of (or alongside) the static auto_remediate input, for a human-reviewed revocation on production clusters.
  • Pair this with a separate, manually-triggered flow that re-enables login (ALTER ROLE ... WITH LOGIN = true) for a specific role during a planned maintenance window.

Pitfalls

  • A role cannot alter the SUPERUSER/LOGIN status of the role it is currently authenticated as. Verified from Cassandra's own documentation: "A role with SUPERUSER status can alter the SUPERUSER status of another role, but not the role currently held." If cassandra_username (or the session's authenticated role) is itself among the flagged rows, revoke_role_login will fail for that specific row even with both gates open; this must be remediated by a different superuser role.
  • This audits direct superuser flags only, not inherited privilege. member_of role membership can grant superuser-equivalent access transitively; verified from system_auth.roles's own documented columns, this blueprint's AT_RISK check only inspects each row's own is_superuser/can_login, not its membership chain.
  • allowFailure: true on query_superuser_roles is what lets UNREACHABLE be classified instead of failing the execution outright. Matching this repo's own cassandra-replication-audit-gate.yaml pattern: a connection or authentication failure produces an empty/undefined rows output rather than aborting the flow, which classify_superuser_exposure reads as UNREACHABLE.
  • auto_remediate and dry_run are independent gates, not a single boolean. Both must be "true"/"false" respectively for the revocation Loop to run; setting only one leaves the other still blocking.
  • Switch case keys are quoted strings matching the classifier's exact output. "HEALTHY", "AT_RISK", and "UNREACHABLE" must match classify_superuser_exposure.value verbatim; the defaults: block exists specifically to catch any value that does not match one of those three cases.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.