Feature selection
Feature selection is the real model choice
When an investment model fails out-of-sample, the algorithm is rarely the culprit. The capacity of modern learning methods is more than enough to fit a cross-section of returns; the question is which columns were allowed to reach the training set. Feature selection — the step that decides what the model is allowed to see — is where most of the real modelling decisions are made, and most of the real mistakes.
The three families of selection method answer different questions. Filter methods — mutual information, rank correlation, simple univariate tests — score each candidate independently of the model and keep the top scorers. They are fast and transparent, and they make no claim about how features interact. They are also blind to redundancy: two near-identical columns both score well and both survive, doubling the exposure to the same signal.
Wrapper methods — forward selection, backward elimination, recursive feature elimination — wrap selection around the model itself, scoring subsets by the model’s own performance. They catch interactions that filters miss, but they are expensive, and because they use the model’s score on the same data that selects the features, they are generous to themselves. A wrapper run on the full sample is a description of the sample, not a method.
Embedded methods — LASSO, elastic net, stability selection, tree importance — fold selection into the fitting step. LASSO’s L1 penalty drives some coefficients to zero, which is selection as a side effect of regularisation. The penalty strength is itself a choice, usually set by cross-validation, and the set of features that survives can be unstable under small changes in the data.
The algorithm is the loud part of the pipeline. The selection step is the quiet part that decides the result.
The discipline that matters most is not the choice of method but the point-in-time, out-of-sample treatment of the selection itself. If features are selected on the full sample and the model is then “tested” on the same sample, the test measures nothing. If features are selected on a training window and frozen for the test window, the test measures whether the selected set generalises — which is the only question that matters for live use. The difference between these two protocols is the difference between a model and a memory.
Stability is the second quiet question. Run a selection method on a hundred bootstrap samples of the same training window; the features that appear in most runs are the ones worth trusting. A feature that appears in three runs out of a hundred is not a discovery; it is noise that happened to align with one sample. Stability selection formalises this by keeping only features that survive across resamples, and it is one of the cheapest ways to turn an overfit selection step into a defensible one.
None of this is resolved by switching to a more powerful model. A gradient-boosted ensemble fed a leaky column will fit the leak beautifully and report a low training error with great confidence. The confidence is real; the leak is real; the future performance is not. Feature selection is not a preprocessing chore to be finished before the interesting work begins. It is the interesting work, and it is where an investment model is, quietly, decided.