Download icon
Script icon
Parallel icon
Queries icon
If icon
Log icon
Pause icon
SlackIncomingWebhook icon
Schedule icon

Backtest Overfitting Guard with Deflated Sharpe Ratio and PBO

Test whether a backtest's best strategy beats luck with the Deflated Sharpe Ratio and PBO, using Kestra, Python and DuckDB on free Fama-French data.

Categories
BusinessData

Trying many strategy variants and reporting the best one is how most backtests end up overfit: with enough variants, one of them looks good by chance. This blueprint runs a grid of long-short industry momentum strategies on 25 years of daily data, then asks whether the winner is distinguishable from luck using two tests from the quantitative finance literature. A result that fails either test is held for a researcher instead of being published.

The demo runs with no API keys or accounts: the data comes from Kenneth French's public data library. With the defaults, the best variant has an annualized Sharpe ratio of about 0.27, almost exactly what the best of 12 unskilled strategies would be expected to reach, so the gate pauses the execution for review.

How it works

  1. download_industry_returns (io.kestra.plugin.core.http.Download) fetches daily returns of the 49 Fama-French industry portfolios.
  2. run_strategy_grid (io.kestra.plugin.scripts.python.Script) backtests every combination of lookback_days and holding_days. Each variant ranks industries on trailing return up to day t, goes long the top_n best and short the top_n worst from day t+1, and rebalances every holding period. Daily returns per variant are written to trial_returns.csv.
  3. evaluate (io.kestra.plugin.core.flow.Parallel) runs two tasks on that file side by side:
    • leaderboard (io.kestra.plugin.jdbc.duckdb.Queries) unpivots the variants and computes annualized return, volatility, Sharpe ratio and maximum drawdown with window functions.
    • overfitting_tests (io.kestra.plugin.scripts.python.Script) computes the Deflated Sharpe Ratio (DSR) of the best variant, which corrects its Sharpe ratio for the number of variants tried and for non-normal returns, and the Probability of Backtest Overfitting (PBO) through combinatorially symmetric cross-validation: across every split of the sample into halves, how often the in-sample winner ranks in the bottom half out of sample.
  4. overfitting_gate (io.kestra.plugin.core.flow.If) checks DSR against dsr_threshold and PBO against pbo_threshold. A result that fails either check is logged as a warning and researcher_review (io.kestra.plugin.core.flow.Pause) holds the execution for up to seven days. Resuming it publishes the report; killing it discards the result; with no decision, it fails.
  5. build_report (io.kestra.plugin.scripts.python.Script) writes report.md with the verdict, the test statistics and the full leaderboard.
  6. notify (io.kestra.plugin.core.flow.If) posts the verdict to Slack when notify_slack is true.

What you get

  • A reproducible research pipeline that refuses to publish a result it cannot distinguish from luck.
  • {{ outputs.build_report.outputFiles['report.md'] }}: a Markdown note with the verdict and leaderboard.
  • {{ outputs.leaderboard.outputFiles.leaderboard }}: the leaderboard as CSV.
  • {{ outputs.run_strategy_grid.outputFiles['trial_returns.csv'] }}: daily returns of every variant, ready for further analysis.
  • {{ outputs.overfitting_tests.vars }}: best_trial, best_sharpe_annualized, expected_max_sharpe_annualized, deflated_sharpe_ratio, pbo and cscv_combinations.

Who it's for

Quant researchers, data scientists and students who run parameter sweeps and want an objective check before a backtest reaches a paper, a deck or a production allocation.

Why orchestrate this with Kestra

  • The review gate is part of the pipeline, not a convention: a flagged result cannot be published without someone resuming the execution, and every decision is recorded in the execution history.
  • Each step's inputs, outputs and files are stored per execution, so any reported number can be traced back to the exact data and parameters that produced it.
  • Python, DuckDB and notifications are declared in one file, and the schedule reruns the study as new data is published.

Prerequisites

  • Docker available to the Kestra worker. The Python tasks run in the default python container and install kestra, numpy and pandas with uv.
  • Outbound HTTPS to mba.tuck.dartmouth.edu (Kenneth French data library) and to PyPI.
  • Optional, only when notify_slack is true: the secret SLACK_WEBHOOK_URL, a Slack incoming webhook URL used to post the verdict.

Inputs

Input Type Default Purpose
start_year INT 2000 First year of the sample. Industries with missing data after it are dropped.
lookback_days ARRAY of INT [21, 63, 126, 252] Momentum formation windows, in trading days.
holding_days ARRAY of INT [5, 21, 63] Days between rebalances. The grid has one variant per lookback and holding pair.
top_n INT 5 Industries held long and short.
cscv_blocks INT 16 Even number of blocks for CSCV. 16 blocks give 12,870 splits.
dsr_threshold FLOAT 0.95 Minimum DSR to publish without review.
pbo_threshold FLOAT 0.5 Maximum PBO to publish without review.
notify_slack BOOL false Post the verdict to Slack.

Quick start

  1. Execute the flow with the defaults. It takes about a minute and a half on a laptop, most of it installing Python dependencies on the first run.
  2. The execution pauses at researcher_review. Open the Logs tab to read the warning with the DSR and PBO values, then resume the execution from the UI to publish the report.
  3. Open the Outputs tab and preview report.md on build_report.
  4. To see the passing path, execute again with dsr_threshold set to 0.5.
  5. Enable the weekly_after_data_refresh trigger once the defaults suit your research.

How to extend

  • Replace run_strategy_grid with your own strategy: any task that writes a trial_returns.csv with a date column and one column of daily returns per variant works with the rest of the flow.
  • Use a different dataset from the same library, such as the 10 or 25 size and book-to-market portfolios, by changing the uri and the table name the parser looks for.
  • Add an onResume input to researcher_review to record the reviewer's decision and comment.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.