Model evaluation

Leakage: the quiet bias that inflates every backtest

A backtest that looks too good is almost never too good. It is usually the result of a column that knew the answer — a feature computed with information that would not have been available on the day the trade was meant to be placed. This is data leakage, and it is the single most common reason that an investment model publishes a handsome backtest and then loses money the moment it meets a live order book.

Leakage arrives in three quiet forms. The first is lookahead bias: a feature that uses a value dated after the prediction. A ratio that uses the full-year reported earnings, available only months after the fiscal year ends, cannot be known on the first of January; using it as a January feature is using the future. The fix is a strict point-in-time index: every column carries the date on which its value first became knowable, and a prediction at time t may only see columns dated t or earlier.

The second is survivorship bias: a universe that contains only the firms that survived to the end of the sample. A backtest on today’s index members, run over the last twenty years, excludes every firm that was delisted, bankrupted or merged away in between — precisely the firms whose returns would have hurt. The average return of the survivors is not the average return of the universe; it is the average return of the winners, and it is always too high.

The third is label leakage: a target that shares information with a feature. If the target is next-month return and a feature is a twelve-month return that ends next month, the feature contains the target. The model learns the trivial identity and reports it as alpha. This is the hardest to catch because nothing in the data looks wrong; the wrongness is in the overlap of the windows.

A feature that predicts is not the same as a feature that remembers.

The discipline that holds all three in check is the same one that governs feature selection: a point-in-time, walk-forward protocol in which every column is timestamped, the universe is reconstructed as it stood on each date, and the model is retrained only on data that was available at each decision point. It is slower than a vectorised backtest, and it is the only kind of backtest whose numbers describe the future rather than the past.

A practical check is to ask, of every feature, the question that leakage hides: on what exact date did this value first become available? If the answer is “the date stamped on the column”, the feature is honest. If the answer is “later than that”, the feature is leaking, and no amount of cross-validation will repair it — cross-validation on a leaking column produces confident, well-validated leakage.

The uncomfortable part is that leakage is rarely malicious. It is usually the default behaviour of a tidy dataset that has been aligned to the end of the sample for convenience. Treating every dataset as leaking until it proves otherwise is not cynicism; it is the minimum standard for a backtest worth reading.

Back to all articles