Correlation Is Not Causation: The Five Reasons Why

September 9, 2026

Statistics · Relationships
Correlation Is Not Causation: The Five Reasons Why
Everyone can recite the slogan. Almost nobody can list the actual reasons two things might correlate without one causing the other. Here are all five.
“Correlation isn’t causation” is the most repeated phrase in statistics and one of the least understood. People say it, nod wisely, and then quietly go on assuming that a strong correlation basically means causation anyway.
So let’s do better than the slogan. If X and Y move together, there are exactly five things that could be going on — and only one of them is “X causes Y.” Learn all five and you’ll never be fooled by a correlation again, and you’ll be able to say precisely why a claim doesn’t hold.

The five explanations for any correlation

Suppose ice cream sales (X) and drowning deaths (Y) rise and fall together. Why might that be?

1. X causes Y

The one everyone jumps to. Sometimes it’s right! But it’s only one of five possibilities, and jumping to it is the whole mistake. (Ice cream does not cause drowning.)

2. Y causes X (reverse causation)

Maybe you’ve got the arrow backwards. “Police numbers correlate with crime” — does more policing cause crime, or does more crime cause cities to hire police? The direction isn’t always obvious, and correlation is perfectly symmetric: it can’t tell X→Y from Y→X.

3. A third variable causes both (confounding)

This is the ice cream one. Hot weather drives up both ice cream sales and swimming, which drives up drownings. Neither causes the other; a lurking third variable causes both. This is the single most common reason correlations mislead.
🔑 Key term
Confounder — a variable that influences both X and Y, creating a correlation between them even though neither causes the other. Confounding is why observational data can’t easily prove causation: there’s always another lurking variable you haven’t ruled out.

4. Coincidence (spurious correlation)

Sometimes there’s no connection at all — the correlation is a fluke. Test enough pairs of unrelated things and some will line up by chance. US cheese consumption correlates almost perfectly with the number of people who died tangled in their bedsheets. Nobody thinks these are related; they just happened to trend together.

5. Selection / bias in how the data was collected

The correlation can be an artefact of who ended up in your sample. If you only study hospitalised patients, or only successful firms, the very act of selection can manufacture a correlation that doesn’t exist in the wider world.
💡 Insight — the burden of proof
To claim “X causes Y,” you have to rule out the other four. That’s hard — which is exactly why the gold standard for causation is the randomised controlled experiment. Randomly assigning who gets X breaks the link between X and any confounder, kills reverse causation (X now comes first, by design), and neutralises selection. When you can’t randomise, you’re left arguing that you’ve somehow ruled out reasons 2–5 — which is the entire discipline of causal inference.
📊 Case study — when getting this wrong costs lives
For decades, observational studies showed that women taking hormone replacement therapy (HRT) had lower rates of heart disease. The correlation was strong and consistent. Doctors concluded HRT protected the heart and prescribed it widely.
Then, in 2002, a large randomised trial — the Women’s Health Initiative — tested it properly. HRT did not protect the heart. In some groups it slightly raised cardiovascular risk.
What had happened? Confounding. Women who took HRT tended to be wealthier, healthier, and more health-conscious to begin with. Their lower heart disease came from those advantages, not from the drug. The third variable — socioeconomic status and baseline health — caused both “takes HRT” and “healthier heart.”
This wasn’t a minor academic error. Millions of prescriptions rested on a correlation that a randomised experiment overturned. “Correlation isn’t causation” is not a pedantic slogan — it’s a lesson written in real medical harm.
⚠ Common error — “but we controlled for confounders”
Controlling for variables in a regression handles the confounders you measured and thought of. It does nothing for the ones you didn’t — and the dangerous confounder is usually the one nobody measured. This is why observational studies with lots of controls still can’t match a randomised experiment, and why the honest write-up says “associated with,” not “causes.” For how this plays out in regression specifically, see omitted variable bias.

Practice questions

Q1. Countries with more Nobel laureates per capita consume more chocolate. Name the most likely explanation from the five.
Q2. Students who attend more lectures get higher grades. Give one reason this might not be “attendance causes grades.”
Q3. Among hospitalised patients, being a smoker is associated with lower mortality from a certain disease. Why might this be a selection artefact rather than smoking being protective?
Q4. Sales of a product and the marketing spend on it correlate strongly. Give the reverse-causation story.

Worked answers

A1. Confounding (reason 3). National wealth plausibly drives both — richer countries can afford more chocolate and more world-class research institutions. Chocolate doesn’t make Nobel winners; a lurking third variable inflates both.
A2. Several work: confounding — conscientious, motivated students both attend more and study harder, so motivation drives both. Or selection — struggling students may drop out, leaving a biased sample. The point is you can’t conclude that forcing a random student to attend more would raise their grade.
A3. This is a selection artefact (reason 5, a collider). You’ve conditioned on being hospitalised. Non-smokers who are hospitalised for this disease may tend to have more severe underlying illness (since they lack the “obvious” smoking risk factor, something worse got them there), so within the hospital, smokers look healthier. In the general population no such protection exists — the correlation is created by studying only the hospitalised.
A4. Reverse causation (reason 2). Many firms set marketing budgets as a percentage of expected or recent sales. So high sales cause high marketing spend, not the other way around — or at least the arrow runs both ways, and a naive correlation can’t separate them.

The short version

• A correlation has five possible explanations, only one of which is “X causes Y.”
• The other four: reverse causation, confounding, coincidence, selection.
Confounding is the most common trap — a third variable driving both.
• Controlling for confounders only handles the ones you measured.
• The gold standard for causation is the randomised experiment, which rules out the other four by design.

References

1. Writing Group for the Women’s Health Initiative (2002) “Risks and Benefits of Estrogen Plus Progestin in Healthy Postmenopausal Women,” JAMA, 288(3), pp. 321–333.
2. Pearl, J. & Mackenzie, D. (2018) The Book of Why. New York: Basic Books.
Causation is the thread running through Statistics Made Simple.
From confounders and colliders to natural experiments — the tools for arguing causation when you can’t randomise.

Related Posts

Marginal Analysis: Why Every Marginal Concept Is a Derivative

Marginal cost, marginal revenue, marginal utility, marginal product – four names for one piece of mathematics. Why marginal always means derivative, why fixed costs vanish at the margin, why MC cuts AC at its minimum, and why MR = MC is simply the profit derivative set to zero.

t-test vs z-test: Which One Do You Actually Use?

One question decides it: do you know the population standard deviation? If yes, z-test; if you’re estimating it from your sample, t-test — which is almost always. Why the t-distribution has fatter tails, the Guinness brewery origin story, and why “n > 30 use z” is a shortcut, not the rule.

Simpson’s Paradox Explained: When Aggregate Data Lies to You

Treatment A beats Treatment B in every single subgroup of patients — and yet, combine all the subgroups together, and Treatment B wins overall. This isn't a contradiction or a calculation error. It's a real, well-documented phenomenon called Simpson's Paradox, and the...