Hidden Markov Model: What Market Regime Labels Mean

You open a backtest and one line stops you. The strategy, it reports, only performs in the low-volatility regime. Somewhere a stretch of the chart has been tagged calm and another stretch tagged turbulent, and the whole result rests on that split. So where did the label come from, and how much weight should it carry before you let it change how you read the chart?

Often the answer is a regime model, and one of the most common is the hidden Markov model, or HMM. It’s worth understanding what that label actually represents before it shapes a research conclusion or a trading decision. A regime label is a modelled summary of the data, produced under assumptions you can inspect. Read it that way and it earns its place. Read it as a fact the market handed you and it will quietly mislead you.

What a hidden Markov model actually estimates

Start with the split between what you can see and what you can’t. The observable series is the data you record directly: daily returns, a realized volatility measure, traded volume. The hidden state is the thing you never observe, the condition the market is supposedly in on a given day. An HMM assumes each observation is drawn from a distribution that depends on which hidden state is active, and it works backward from the numbers you can see to a probability for the state you can’t.

Two states are enough to make this concrete. Call them relatively calm and relatively turbulent. When I fit a two-state model to daily returns, the calm state usually carries a much smaller daily standard deviation than the turbulent one, often on the order of 0.7 percent against 2 percent or more, with the calm state’s returns clustered nearer zero. The model has no idea what the words calm and turbulent mean. It only sees that the data appear to come from two different distributions, and it estimates the parameters of each. The names are ours, pinned on afterward.

The three pieces every regime model carries

Every HMM rests on three components, and knowing them tells you what the estimate can and cannot do.

The first is state-specific behaviour: what the data tend to look like inside each state. In the two-state case that’s two sets of parameters, one describing the calm distribution and one the turbulent distribution. The second is the set of transition probabilities, the chances of moving from one state to another between one day and the next. Picture a small two-by-two grid: the top row holds the odds of staying calm versus flipping to turbulent, the bottom row the odds of the reverse. A model might estimate a 0.95 probability of staying calm tomorrow given calm today, and only a 0.05 probability of flipping. The third component is the estimation step: given the observations up to a date, the procedure computes which state is most consistent with the data, expressed as a probability rather than a hard yes or no.

That last point does real work. The output assigns a probability to each state: given everything observed so far, the turbulent state carries, say, 0.68. A trader using this framework might treat that number as one input among several, rather than a switch that turns a strategy on or off. Feed the model volume alongside returns and the same machinery runs, though the states it settles on can shift, which is the first hint that the inputs decide a lot.

Why persistence keeps this from being a volatility line

Here’s where an HMM parts company with the simplest alternative, drawing a threshold on a volatility chart. Suppose you decide that any day with 20-day realized volatility above 25 percent annualized counts as turbulent. That rule flips state every single time the measure pokes across the line, even for one day, even on noise. You get a jittery label that changes constantly around the boundary and tells you very little.

The transition probabilities build in persistence, or stickiness. Because the model assigns a high chance of staying in the current state, a single unusual observation doesn’t drag the estimate across on its own. It takes a run of evidence, several days pulling the same way, to shift the probability meaningfully. That’s closer to how conditions actually behave: turbulence tends to arrive and linger rather than switch on for a day and vanish. The stickiness is a modelling choice though, baked into those transition numbers, and set it too high and the model turns slow, acknowledging a genuine change only well after it started. You meet the same tension in a structural break test, where sensitivity trades off against false alarms.

A state label is a modelled summary

This is the part most worth internalising. You set the number of states before fitting. Hand the same data three states and it returns three, each with different parameters and different boundaries. The states are a description fitted to the data under your assumptions, not a hidden truth the model uncovered.

So the label is a summary, useful for compressing a messy series into a handful of conditions you can reason about. It carries no promise that a clean regime exists out there waiting to be found, and it isn’t a forecast of tomorrow. When a regime-factor screening result sorts periods into buckets, the buckets are only as meaningful as the model that drew them. The honest reading treats the label as shorthand for how the data behaved, subject to the choices that produced it.

The label can change after the date it describes

A regime estimate isn’t fixed once you compute it. Run the model using only the data available up to a given day and you get one probability for that day, the filtered estimate. Run it again later, letting it use the observations that came afterward, and the probability for that same day can move. This second version, the smoothed estimate, is the one you often see plotted in a finished research chart, and it quietly uses information that wasn’t available on the date it labels.

I’ve watched a smoothed label repaint a stretch that looked calm in real time into a turbulent state once the following weeks were folded in. Nothing about the past had changed. The model simply had more to work with. For anyone reading a backtest, that’s the trap: a strategy that looks like it cleanly sidestepped turbulent periods may only look that way because the labels were drawn with hindsight. This is the concern behind look-ahead bias, and it’s why the timing of the information matters as much as the label itself.

How to read a regime claim before you trust it

When someone hands you a chart with the periods already coloured in, or a research result that leans on a regime split, a few questions separate a summary you can use from one you can’t. I run through the same short list every time.

  • What inputs went in? Returns alone, or volatility and volume too? The choice of series shapes every state the model finds.
  • How many states were imposed? Two, three, more? That number was a decision, not a discovery, and it fixes the shape of the answer.
  • Do the labels hold up across samples? Refit on a different window. If the states and boundaries move around, the split is fragile.
  • Does the analysis respect the information available on each date? Filtered estimates that only use the past are honest. Smoothed ones drawn with hindsight flatter a backtest.

That last question is where a walk-forward analysis does its job. By only ever using data available on each historical date, it tests whether a regime signal would actually have helped in real time, rather than whether it looks tidy in retrospect. Regime detection is still an open research question, with new methods proposed regularly, so a fresh technique is a reason for interest, not proof that any single approach works broadly.

Read the label, then read past it

A hidden Markov model gives you a common language. It lets you talk about why a relationship looks different across periods without pretending you can pin the switch precisely in real time. That’s genuinely useful, and it’s a cleaner way to reason about changing conditions than eyeballing a volatility chart and calling it.

Hold on to the limitation just as tightly. State counts, input choices, sample periods, and estimation methods can each move the result materially, and two careful analysts can label the same history differently and both be defensible. Markets can also shift in ways no fixed number of states was built to capture, which is the deeper point in Nassim Taleb’s writing on fat tails: the rare, structure-breaking move is exactly the one a tidy model tends to underweight. So use the label as a lens, question how it was made, and never let it stand in for the chart in front of you. Learn the pattern. Ride the trend. Keep the gains.

Educational content only. Not investment advice. Trading involves risk. You are responsible for your decisions.

Get the free Market Wisdom e-book

Join Trends and Breakouts — historical winners, breakout studies, and risk lessons. No spam, unsubscribe anytime.