Script icon
UploadFiles icon
Switch icon
Set icon
Log icon
Fail icon
Schedule icon

Read out an A/B test with a sample ratio check and a guardrail metric

Read out A/B tests in Kestra. Detect sample ratio mismatch, test conversion with a confidence interval, check a guardrail, record the decision.

Categories
BusinessData

Most bad experiment decisions come from reading the result in the wrong order. A treatment that looks like a winner can be the product of a broken split: a logging bug drops some treatment visitors who did not convert, so the conversion rate rises and the test reports a significant lift that does not exist. This flow reads the experiment out in the order that prevents that:

  1. Sample ratio mismatch (SRM). A chi-square test of the observed split against the planned 50/50. A p-value under srm_alpha means the experiment is invalid, and no other number is trusted.
  2. Primary metric. Conversion, with a two-proportion z-test and a 95% confidence interval on the difference.
  3. Guardrail. Refunds per conversion. A significant conversion win that raises refunds by more than max_guardrail_increase is not shipped.

It runs with no setup and no network: the statistics use the Python standard library. The scenario input generates 120,000 visitors for four realistic outcomes.

How it works

  1. readout (io.kestra.plugin.scripts.python.Script, python:3.12-slim, no packages) generates the events, computes the three checks, writes readout.md and outputs the numbers and a verdict.
  2. publish_readout (io.kestra.plugin.core.namespace.UploadFiles) writes experiments/<experiment>/readout.md.
  3. decision (io.kestra.plugin.core.flow.Switch):
    • SHIP: record_ship (io.kestra.plugin.core.kv.Set) stores the decision for a feature flag service or a deploy pipeline.
    • INCONCLUSIVE: log the interval and keep running.
    • INVALID: fail and say the split is broken.
    • GUARDRAIL_BREACH and ROLL_BACK: fail with the numbers.
  4. daily (io.kestra.plugin.core.trigger.Schedule).

Quick start

scenario Split (control / treatment), SRM p Conversion Refunds per conversion Verdict
WINNER 59,994 / 60,006, p = 0.97 10.21% to 10.78%, p = 0.0011, CI +0.23 to +0.92 pts 5.03% to 4.84% SHIP
FLAT p = 0.95 10.19% to 9.98%, p = 0.23 flat INCONCLUSIVE
GUARDRAIL_HIT p = 0.97 +5.7%, p = 0.0011 5.03% to 7.64% (+52%) GUARDRAIL_BREACH
BROKEN_ASSIGNMENT 60,410 / 56,471, p = 1e-30 +6.4%, p = 0.0003 flat INVALID

The last row is the point of the flow. The conversion test alone would ship it as a significant win, even though the treatment has no effect.

Inputs

  • experiment (STRING), scenario (SELECT, demo only).
  • alpha (FLOAT, default 0.05), srm_alpha (FLOAT, default 0.001), max_guardrail_increase (FLOAT, default 0.10).

Expected outputs

  • outputs.readout.vars: verdict, visitors per arm, srm_p, both rates, relative_lift, p_value, ci_low_pts, ci_high_pts, refund rates and refund_increase.
  • Namespace file experiments/<experiment>/readout.md and, on SHIP, KV experiment_<experiment>.

Using it with your data

  • Replace the generator with a query that returns visitors, conversions and guardrail events per arm, for example a io.kestra.plugin.jdbc query before the script.
  • Decide the sample size before the experiment starts and read the result once. Reading a fixed-horizon test every day and stopping on the first significant day inflates false positives.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.