Simpson’s Paradox Explained: When Aggregate Data Lies to You

August 31, 2026
Treatment A beats Treatment B in every single subgroup of patients — and yet, combine all the subgroups together, and Treatment B wins overall. This isn’t a contradiction or a calculation error. It’s a real, well-documented phenomenon called Simpson’s Paradox, and the Statistics Made Simple Complete Bundle works through the exact case study that made it famous.

In 1973, the University of California, Berkeley was sued for gender discrimination in its graduate admissions. The raw numbers looked damning: men were admitted at 44%, women at only 35%. But when statisticians examined admission rates department by department, they found something startling — in most individual departments, women were admitted at equal or higher rates than men. The overall gap wasn’t discrimination in admissions decisions; it emerged because women disproportionately applied to more competitive departments with lower acceptance rates for everyone. This reversal — where a trend appears in several different groups of data but disappears or reverses when the groups are combined — is called Simpson’s Paradox, and it is one of the most important cautionary tales in all of applied statistics.

What Is Simpson’s Paradox?

Simpson’s Paradox occurs when a trend that appears consistently within several separate groups reverses when those groups are aggregated into one. It’s named after statistician Edward Simpson, who described the phenomenon formally in 1951, though earlier statisticians had noted similar effects decades before.

The paradox isn’t a mathematical error or a trick — every individual calculation is correct. What makes it feel paradoxical is that our intuition strongly (and usually reasonably) expects that if A beats B in every subgroup, A must beat B overall. Simpson’s Paradox is the proof that this intuition, while often right, is not guaranteed — and the mechanism behind why it fails is a confounding variable that differs in size across the groups being combined.

Worked Example: The Two Hospitals

Two hospitals, A and B, both perform a risky surgery. Hospital A reports a 90% survival rate; Hospital B reports 85%. On the surface, Hospital A looks better. But look at the breakdown by patient condition:

Hospital A Hospital B
Healthy patients 590/600 survived (98.3%) 190/200 survived (95.0%)
Critical patients 210/400 survived (52.5%) 640/800 survived (80.0%)
Overall 800/1000 survived (80.0%) 830/1000 survived (83.0%)

Hospital B has a higher survival rate than Hospital A for both healthy patients and critical patients individually — 95.0% vs 98.3% is actually reversed here for illustration, let’s recompute correctly: look again — Hospital A does better with healthy patients (98.3% vs 95.0%) but Hospital B does dramatically better with critical patients (80.0% vs 52.5%). Because Hospital B treats far more critical patients as a share of its total caseload (800 out of 1,000, vs Hospital A’s 400 out of 1,000), Hospital B’s overall rate is pulled down by its heavier critical caseload — yet it still ends up with the higher combined rate because its performance advantage in the critical category is so large. The confounding variable here is patient condition, which differs sharply in proportion between the two hospitals and is exactly what a simple headline comparison misses.

Chapter 8 of Statistics Made Simple works through the full Berkeley admissions dataset step by step, showing exactly how department-level acceptance rates combine into the university-wide figures.

Why Does This Happen? The Weighting Mechanism

Simpson’s Paradox arises specifically because the two groups being compared don’t just differ in their within-group rates — they also differ in the relative size of their subgroups. An overall rate is a weighted average of the subgroup rates, weighted by how many observations fall into each subgroup. If Group A’s overall rate is weighted heavily toward a subgroup where it happens to underperform, while Group B’s overall rate is weighted heavily toward a subgroup where it happens to outperform, the subgroup-level pattern can invert once the weights are applied. The paradox is entirely a statement about how weighted averages behave — nothing paradoxical is happening mathematically, only something counterintuitive.

The Berkeley Admissions Case, Explained

Returning to the opening example: at Berkeley, women applied in much greater numbers to departments like English and Sociology, which had lower overall acceptance rates for everyone (both men and women), while men applied disproportionately to departments like Engineering, which had higher acceptance rates for everyone. When you aggregate across all departments, the university-wide numbers make it look like women faced a systematic admissions disadvantage — but the department-by-department numbers showed the opposite pattern in most departments individually. The confounding variable was choice of department, correlated with both gender and with each department’s overall admission difficulty.

This finding, published by Bickel, Hammel and O’Connell in 1975, remains one of the most cited real-world examples of Simpson’s Paradox precisely because it demonstrates how a socially and legally important conclusion can flip entirely depending on whether you look at aggregated or disaggregated data — and because it illustrates that the “obvious” reading of aggregate data can sometimes obscure rather than reveal the true underlying pattern.

Statistics Made Simple
When the data itself can mislead you.
246 pages of explanation and 1,569 practice questions with fully worked answers — Simpson’s Paradox, confounding, and correlation vs causation covered together as a connected set of critical-thinking tools.

Get both books — $16, save $5 →

How to Detect and Avoid Being Misled

Always ask what’s being averaged over. Whenever a comparison spans multiple natural subgroups — departments, hospitals, regions, time periods — check whether the subgroup composition differs between the two things being compared. If it does, the aggregate comparison may not mean what it appears to mean. Look for a lurking confounding variable that could plausibly differ in proportion between the groups — patient severity, department competitiveness, company size — anything correlated with both the grouping variable and the outcome. Report disaggregated results alongside aggregates whenever subgroup composition is uneven, since the disaggregated view is often the one that reflects the real underlying relationship, while the aggregate can be an artefact of how the groups happened to be weighted.

Real-World Applications

Simpson’s Paradox appears constantly outside the classroom, often with serious stakes. In clinical trials, a drug can appear to help patients overall while a closer look by age group, disease severity, or comorbidity reveals it actually helps some subgroups and harms others — which is precisely why modern trials pre-register subgroup analyses rather than relying solely on the topline result. In batting averages in baseball, a famous example shows one player having a higher batting average than another in both of two separate seasons, yet a lower combined average across both seasons — purely because of how many at-bats occurred in each season. In economic policy evaluation, a minimum wage increase might appear to reduce employment in aggregate data while actually having no effect (or even a positive effect) within every individual industry, if industries with naturally declining employment happened to be more affected by the policy change for unrelated reasons.

Simpson’s Paradox vs Ordinary Confounding

Simpson’s Paradox is a specific, extreme case of the more general confounding problem discussed in relation to correlation and causation elsewhere in this series. Ordinary confounding can bias a relationship without actually reversing its direction. Simpson’s Paradox is the special, more dramatic case where the confounding is strong enough, and the subgroup weights different enough, that the direction of the relationship flips entirely between the disaggregated and aggregated views — making it an especially memorable illustration of why “controlling for” the right variables matters so much in observational research.

Common Mistakes

Assuming the aggregate result is always the “true” one. There’s no universal rule that combined data is more trustworthy than subgroup data — in Simpson’s Paradox cases, it’s frequently the disaggregated, subgroup-level view that reflects the real relationship, while the aggregate is distorted by uneven weighting.

Assuming the disaggregated result is always the “true” one either. The correct interpretation depends on the actual causal structure of the situation, not a blanket rule in either direction — sometimes the aggregate genuinely is the more meaningful number, and careful subject-matter reasoning about the confounder is required, not just a mechanical preference for one level of aggregation over the other.

Failing to check subgroup sizes before trusting a comparison. Whenever comparing rates or averages across groups that themselves contain meaningfully different natural subgroups, it’s worth explicitly checking whether the subgroup composition is similar between the groups being compared — if it isn’t, Simpson’s Paradox is a live possibility.

Practice Question

Two salespeople are compared. In Q1, Alex closed 8 of 10 leads (80%) and Sam closed 45 of 60 leads (75%). In Q2, Alex closed 2 of 40 leads (5%) and Sam closed 3 of 70 leads (about 4.3%). Alex’s manager claims Alex outperformed Sam both quarters. Check the combined totals and explain what’s happening.

Answer: Combined, Alex closed (8+2)/(10+40) = 10/50 = 20%. Sam closed (45+3)/(60+70) = 48/130 ≈ 36.9%.

Despite Alex having a (very slightly) higher close rate in both individual quarters, Sam has a substantially higher combined close rate. This is Simpson’s Paradox: Alex handled far more leads in Q2, the much harder quarter (5% close rate for everyone) relative to Q1, while Sam handled a more even split with more leads concentrated in the easier Q1 period. The confounding variable is lead difficulty by quarter, and it’s weighted very differently between the two salespeople’s caseloads — making the combined comparison misleading without the quarter-by-quarter breakdown.

Spotting a Simpson’s Paradox setup in an exam question takes pattern recognition built through repetition. The Statistics Made Simple Practice Questions workbook has a dedicated section of these scenario-based problems with full model answers.

References

1. Simpson, E.H. (1951) ‘The Interpretation of Interaction in Contingency Tables’, Journal of the Royal Statistical Society, Series B, 13(2), pp. 238–241.
2. Bickel, P.J., Hammel, E.A. and O’Connell, J.W. (1975) ‘Sex Bias in Graduate Admissions: Data from Berkeley’, Science, 187(4175), pp. 398–404.
3. Pearl, J. and Mackenzie, D. (2018) The Book of Why. Basic Books.
4. Moore, D.S., McCabe, G.P. and Craig, B.A. (2021) Introduction to the Practice of Statistics. W.H. Freeman.
5. Wagner, C.H. (1982) ‘Simpson’s Paradox in Real Life’, The American Statistician, 36(1), pp. 46–48.

Related Posts

Outlier Detection: The IQR Method, Z-Scores and What to Do Next

A single data-entry typo — a salary of $5,000,000 instead of $50,000 — can single-handedly move a mean, inflate a standard deviation, and quietly wreck a regression model. The Statistics Made Simple Complete Bundle covers exactly how to catch outliers like this before...