Someone hands you a backtest. The headline reads: an information coefficient of 0.06, a smooth equity curve, and a claim that the model ranks stocks better than the market does. The instinct is to trust the figure, because 0.06 looks precise and precision reads as truth. The harder question is what that number was measured against, and whether it holds up once you trade the way you actually trade.
The information coefficient, usually shortened to IC, is one of the most quoted and least interrogated statistics in quantitative screening. I’ve watched a good-looking IC lose most of its shine the moment I asked which timestamp the scores were taken from. This guide covers what the number measures, the design choices that quietly set its meaning, and a checklist you can run before any ranking metric earns a vote in a decision.
What the information coefficient actually measures
An IC is a correlation between a model’s scores at one moment and the realized outcomes over a defined forward window, across a defined set of names. Rank 500 stocks on some score today. Twenty trading days later, measure each stock’s return over that window. Compute the correlation between the score and the forward return. That single correlation is the IC for that period.
The value runs from minus 1 to plus 1. A positive IC says higher-scored names tended to do better over the window, and a negative IC says the ordering worked in reverse. Magnitude tells you how tightly the order held. In cross-sectional equity work, many practitioners treat a monthly IC around 0.02 to 0.05 as a real, tradeable signal, and treat sustained readings above 0.10 as a reason to check for a mistake rather than a reason to celebrate. The number is small on purpose. You’re trying to order hundreds of noisy assets, so even a faint, repeatable tilt is worth something.
Why ranking models lean on a rank correlation
A ranking model has one job: put the names it expects to do better above the names it expects to do worse. It isn’t asked to predict that a stock will return exactly 7.3 percent, and that distinction shapes how you should score the model. If order is what matters, then the honest measure is a rank correlation, the Spearman version, which asks whether higher-ranked names tended to post higher forward returns without caring about the exact size of each move.
Here the choice between Pearson and Spearman does real work. A Pearson IC computed on raw returns can be dominated by one or two extreme movers. One name up 300 percent can carry the entire figure, so the model looks skilled when it simply caught a single tail. A rank correlation caps each name’s influence at its position in the line, which is why cross-sectional research usually reports it. This is the same lens behind cross-sectional versus time-series momentum, where the question is relative order across a universe rather than one asset’s trend through time. Ranking stocks against each other is an old craft. William O’Neil built relative strength ranking into a discipline long before anyone printed an IC to grade it.
One period, many periods, and the IC ratio
A single-period IC is noisy enough to mislead you on its own. This month’s 0.09 can be next month’s minus 0.04, and neither reading means much in isolation. The average IC measured across many periods is the more honest summary, since it asks whether the ordering held up on repetition rather than on one lucky window.
An average still hides how much the number moved around. That’s where the information coefficient ratio, the ICIR, comes in. Take the mean of the period ICs and divide it by their standard deviation. A mean IC of 0.03 against a standard deviation of 0.06 gives an ICIR of 0.5. In my own screening notes I treat a monthly mean IC under 0.02 as noise until the ICIR earns a second look. A higher ratio means the ordering was more consistent from period to period, not merely positive on average. Treat both figures as descriptive summaries of history. A mean IC of 0.04 sitting on a standard deviation of 0.12 flips sign often, and an edge that flips sign through most of the quarters you’d hold it isn’t one you can lean on.
The design choices that set the meaning
Most of an IC’s meaning lives in decisions made before the correlation is ever computed. The eligible universe, the score timestamp, the return horizon, the treatment of delistings and corporate actions, missing data, ties, sector effects, rebalance timing, and whether observations overlap all move the number, sometimes more than the model does.
Start with the timestamp, because it’s the quietest trap. If the score uses any data that wasn’t knowable at the moment it claims to rank, the IC is inflated by information from the future. That’s textbook look-ahead bias, and it flatters almost any metric it touches. The forward returns matter just as much. They have to run on a properly adjusted price series, with splits and dividends handled correctly, or the outcome you’re ranking against is wrong before the math even starts.
Delistings are where a clean-looking IC often hides its worst sin. If the test quietly drops names that were delisted, halted, or acquired at a loss, it removes many of the outcomes the model should have been penalized for missing, and the IC on the survivors overstates the real edge. That’s survivorship bias wearing a statistic. Overlap is subtler still. Measure 60-day forward returns every single day and consecutive samples share most of their bars, so the effective number of independent observations is far smaller than the row count suggests, and any significance test run on the raw rows looks stronger than the evidence deserves. Sector effects round out the list. A score that is really a bet on one sector will print an IC that reflects that sector’s month, not broad ranking skill, until you neutralize by sector and check whether anything survives.
When a positive IC still fails in practice
A positive average association and a workable strategy are different claims, and the gap between them is wide. An edge can sit almost entirely in the names with the widest spreads, so the ordering is real on paper and gone after trading costs. It can be concentrated in the extreme top and bottom deciles while the middle of the ranking is noise, which matters a great deal if you only ever touch the top slice. It can live in small, illiquid names you can’t size into, so the IC describes an edge with no capacity behind it.
Timing is its own hazard. A signal can post a healthy average IC while running negative through exactly the stretch you’d have held it, so the summary statistic and your lived experience disagree. The average is honest about the whole sample and quiet about your particular holding window, and those two truths part ways more often than most write-ups admit. Crowding adds a slower failure. As more capital ranks stocks the same way, the forward association can decay because everyone is acting on it, a reflexive effect where the measurement erodes the thing it measured. None of this shows up in the headline figure. It shows up when a promising screen becomes a live book, which is why multi-factor composite screening tends to test an IC against costs and capacity before it trusts one.
A checklist for any claimed IC
When a model’s ranking metric lands in front of you, slow down and interrogate it before it earns any weight. A few plain questions separate a measured claim from a marketing number:
- What exactly was ranked: which universe, how many names, and on which score?
- Against which later outcome: what return, measured over what horizon?
- Over what dates, and across how many genuinely independent periods?
- Was the test truly out of sample, or was the score tuned on the same data it was later graded against?
- Is the result statistically meaningful once overlap is accounted for, and economically meaningful after costs, sector neutralization, and capacity limits?
That fourth question is the one people dodge most. A model scored on the same window it was fit to will almost always show a flattering IC, which is the whole reason walk-forward analysis exists: rank on data the model has never seen, then measure. When I read a research write-up now, I check the protocol paragraph before I even glance at the reported coefficient. I want to know whether the out-of-sample test used a simple date split or something closer to purged cross-validation, because overlapping observations inflate the apparent significance even when the window is technically held out. The sample rules and the independence assumption decide what the figure can possibly mean. A ranking metric quoted without its protocol is a headline without a story.
Read the protocol before the number
An IC is a measurement of historical association under one stated protocol, and that’s the whole of what it is. It doesn’t tell you to buy or sell anything, and a strong past coefficient carries no promise that the model keeps ordering names once the regime turns. A trader using a ranking model might treat a stable ICIR as one input among several, and still size every position as though the edge could quietly disappear next quarter. That posture costs little when the signal keeps working and saves a lot when it stops. Learn the number, then learn the protocol behind it. Learn the pattern. Ride the trend. Keep the gains.
Educational content only. Not investment advice. Trading involves risk. You are responsible for your decisions.
Get the free Market Wisdom e-book
Join Trends and Breakouts — historical winners, breakout studies, and risk lessons. No spam, unsubscribe anytime.
