You run a factor screen at the end of August, and the ten names at the top look familiar. You run it again four weeks later. The same companies are still near the top, but the order has scrambled. The stock that sat third is now twelfth, and a name that was outside the top twenty has jumped to fourth. Nothing dramatic happened in the market, which leaves a practical question sitting on the desk: did the screen genuinely change, or did the ordering just shift under a set of scores that barely moved?
Rank correlation is the number that answers that. It measures how consistently an ordered list keeps its relative ordering between two observations, whether those two observations are two dates or two versions of the same model. Before you trust a screen’s output, or before you assume it has drifted, this is the diagnostic that tells you how far the ordering actually traveled.
What rank correlation actually measures
Start with what the number is built from. Take a universe of securities, score every one of them, and sort. You’ll get an ordered list: first, second, third, all the way down. Do it again at a second date or with a second model, and you’ve got a second ordered list of the same names. Rank correlation compares the two orderings and returns a single value, usually between minus one and plus one. Plus one means the two lists agree on every position. Zero means the second ordering carries no information about the first. Minus one means the order flipped end to end.
The word ordering is doing the heavy lifting. Rank correlation ignores the size of each score and looks only at position. If a stock moves from the 5th slot to the 6th, that counts as a small disagreement whether its underlying score slipped by a hair or by a mile. That’s the property that makes the number useful for screens, because a screen’s decision is a sorting decision. What matters is where a name lands in the queue, not the raw score that put it there.
Score level and rank are two different numbers
Here’s the distinction that trips people up, and it’s worth slowing down on. A factor score is a level: a momentum reading of 1.51, a value composite of minus 0.3, a quality z-score of 0.8. A rank is a position in a sorted list: 3rd, 40th, 112th. The two move on different clocks.
I reran a nine-factor screen a month apart on the same universe. One name held its composite almost still, 1.48 against 1.51, a change you’d round away. Its rank fell from 3rd to 19th. Nothing was wrong with the stock. Sixteen names sitting just behind it had gained a little more ground, and that was enough to slide it down the queue. The score barely moved. The rank moved a lot.
The reverse happens too, and it’s the more dangerous case. Two names separated by a whisker of score sit on opposite sides of your top-fifty cutoff. A rounding-level change in either one swaps them across the line, so one enters the portfolio and one doesn’t, on a score difference you can’t see with the naked eye. Ranks are most fragile exactly where you care most, around the cutoff where selection happens. A screen built on multi-factor composite screening inherits this fragility from every factor it blends.
Spearman and Kendall, two ways to compare orderings
Two measures show up most often when people put a number on ordering agreement, and neither one is the universal answer. Spearman rank correlation takes the ordinary correlation formula and feeds it the ranks instead of the raw values. Kendall rank correlation counts pairs: for every possible pair of names, it asks whether the two rankings agree on which one is higher, a concordant pair, or disagree, a discordant pair, then nets them out.
They answer slightly different questions. Spearman reacts to how far names travel, so one stock that plunges forty places drags the number down hard, because squared rank differences dominate the formula. Kendall reacts to how often the order flips, and it treats a single large jump more gently than a scattering of small swaps. For the same two lists, Kendall’s tau usually reads lower than Spearman’s rho. That’s a property of the math and doesn’t mean either list is worse.
That last point is a common misread. A Kendall tau of 0.6 and a Spearman rho of 0.8 computed on the identical pair of rankings don’t disagree with each other. They’re two rulers with different spacing. Pick one, understand what it rewards, and stay with it across your comparisons rather than switching to whichever reads higher on a given day.
A repeatable way to run the check
The mechanics are simple enough to write on an index card, and writing them down is what keeps the comparison honest.
- Fix the eligible universe first. Decide which names are in play before you score anything, so the investable universe is identical on both sides. Comparing a ranking of 500 names against a ranking of 480 confounds ordering change with membership change.
- Compute the score at both points: two dates, or two model specifications, or one live model against one candidate revision.
- Rank both lists, using the same tie handling and the same direction, high score to low.
- Compute the rank correlation across the names the two lists share.
- Then do the part most people skip: look at where the disagreement lives. Sort by the change in rank and read the biggest movers by hand.
The single number is a headline. The list of who moved and why is the story. When I switched one momentum lookback from 126 to 63 trading days, the Spearman rho between the old and new rankings came out near 0.72. The headline said the two screens mostly agreed. The by-hand read said something sharper: almost all of the disagreement sat in the bottom third of the list, names I’d never act on anyway, so for practical purposes the two versions were closer than 0.72 suggested.
Reading a high number and a low number
A high rank correlation, say a Spearman rho of 0.95 between two adjacent months, says the ordering is persistent over that interval. The screen is returning roughly the same queue, and turnover from re-ranking alone won’t be large. That’s often what you want from a screen meant to hold names for weeks or months.
A low number has more than one cause, and the whole diagnostic value lies in telling them apart. Low stability can mean the inputs genuinely changed, which is real signal. It can also mean the measurement is noisy, that the universe shifted underneath you, or that too many names are clustered near a cutoff where tiny score moves produce large rank moves. A screen packed with near-ties around the selection line will print a low rank correlation even in a quiet market, purely from cutoff churn. That last cause is a plumbing problem wearing the costume of market information, and it’s the one worth ruling out first.
This is where a ranked screen meets a long tradition of trading by relative position. Ranking a universe by strength and focusing on the leaders is the spine of William O’Neil‘s approach, and the same fragility applies: the gap between the 40th and 60th strongest name is often statistical noise, even when the ranking looks precise.
Why a stable ranking can still be useless
Now the trap. A stable ranking and a useful ranking are different properties, and a screen can have one without the other. Rank correlation measures agreement between two orderings. It says nothing about whether either ordering predicts returns.
A screen can be perfectly stable and completely uninformative. Rank every stock by the alphabetical order of its ticker and the ranking will hold still month after month, a rank correlation of 1.0, and it’ll forecast nothing. Stability alone only tells you the ordering holds together over time. Whether that ordering picks winners is a separate measurement, and the tool for it is the information coefficient, which correlates the rank against the return that followed.
The other half of the trap runs the opposite way. A genuinely informative signal can, and often should, reorder itself across regimes. A factor that leads in a low-volatility trend can reshuffle hard when the regime turns, and that reshuffle shows up as a low rank correlation between the two periods. Read on its own, the low number looks like a broken screen. Read alongside the regime change, it looks like a signal doing its job. The statistic can’t tell those two stories apart. You’ve got to supply the context.
Ties, gaps, and the housekeeping that moves the number
Plenty of the movement in a rank correlation comes from housekeeping rather than markets, and each piece of plumbing has to be handled the same way on both sides of the comparison.
Ties come first. When several names share a score, the ranking method has to break the tie somehow, and different rules produce different orderings from identical data. Average ranks, first-encountered order, and random tie-breaks each yield a different rank correlation, so the tie rule has to be fixed and identical across both lists.
Missing observations are next. A name with no score on one of the two dates can’t be ranked there, and quietly dropping it changes the universe the correlation is computed over. Corporate actions do the same damage in disguise: a split, a spin-off, or a symbol change can make one date’s data look discontinuous unless the price series is properly adjusted for those events. Liquidity filters that admit a name on one date and exclude it on the next inject membership change straight into the comparison.
Rebalancing frequency sets how often you even ask the question. Compare rankings a day apart and the rank correlation sits high by construction, because little has had time to move. Compare them a quarter apart and the same screen reads far lower. The horizon you choose shapes the number, and it also shapes how much re-ranking your book has to absorb, which ties rank stability directly to portfolio turnover.
Where the number stops talking
Rank correlation earns its place as a first read, and it has a hard edge you should respect. It compares two orderings and reports how much they agree. It doesn’t tell you whether the disagreement mattered.
A month-to-month rho of 0.8 says the ordering shifted a moderate amount. It doesn’t say whether the names that moved were ones you could actually trade at size, whether the reshuffle would have changed your realized return by a basis point or a fortune, or whether any of it holds up outside the particular universe and sample window you measured. Those questions live beyond the summary statistic, in transaction costs, capacity, and out-of-sample testing. The number points you at where to look. It doesn’t do the looking.
So treat rank correlation as a gauge, not a verdict. When a screen’s ordering moves, the value tells you how far and the by-hand read tells you why, and only your own judgment tells you whether the move was worth acting on. Read it that way and it becomes one of the cleaner instruments on the bench for understanding what your screen is really doing from one run to the next.
Learn the pattern. Ride the trend. Keep the gains.
Educational content only. Not investment advice. Trading involves risk. You are responsible for your decisions.
Get the free Market Wisdom e-book
Join Trends and Breakouts — historical winners, breakout studies, and risk lessons. No spam, unsubscribe anytime.
