Game Theory and Nash Equilibrium
When your best move depends on what everyone else does
When your best move depends on what everyone else does
You’ve probably played a game where your move depends entirely on what your opponent does. Chess, poker, even rock-paper-scissors. But here’s something most students don’t realise: economics is full of these situations too.
Should Apple lower its iPhone price? Depends on what Samsung does. Should Nigeria cut oil production? Depends on what Saudi Arabia does. Should a country sign a climate treaty? Depends on whether other countries sign too. These aren’t just business decisions — they’re strategic decisions. And that’s exactly what game theory studies.
Game theory is basically the economics of strategy. It was developed by mathematician John von Neumann and economist Oskar Morgenstern in 1944, then supercharged by John Nash in the 1950s. Nash’s work was so important he won the Nobel Prize in 1994, and his life was turned into the film A Beautiful Mind.
The payoff matrix is just a table showing what each player gets for each combination of choices. Let’s use a real example: two airlines deciding whether to price a route high or low.
| Qatar Airways | ||
|---|---|---|
| High Price | Low Price | |
| Emirates: High Price | 60, 60 | 10, 80 |
| Emirates: Low Price | 80, 10 | 30, 30 ← Nash Equilibrium |
Read it like this: first number = Emirates’ profit (£m), second number = Qatar’s profit (£m).
Now ask yourself: what should Emirates do?
Either way, Low is better for Emirates. That’s a dominant strategy. Qatar faces the exact same situation — Low is always better for them too. So both airlines end up at (Low, Low) = (30, 30). That’s the Nash Equilibrium.
But look at (High, High) — both would get 60 each. They’d be £30m better off! The problem is neither can trust the other to stay there. That instability is the heart of the famous Prisoner’s Dilemma.
Two suspects are caught by police. They’re put in separate rooms and each offered the same deal: rat on your partner or stay silent.
| B: Stay Silent | B: Rat | |
|---|---|---|
| A: Stay Silent | 1 year each | A: 10 years, B: free |
| A: Rat | A: free, B: 10 years | 5 years each ← Nash Equilibrium |
Each person thinks: “If I rat and my partner stays silent, I go free. If I stay silent and my partner rats, I get 10 years. Either way, ratting is better for me.” So both rat — and both get 5 years. Even though if they’d both stayed silent, they’d only get 1 year each.
This is the dilemma: individual logic leads to a collectively worse outcome. Sound familiar? This is exactly why:
John Nash defined equilibrium like this: a set of strategies where no player can do better by switching their strategy, as long as everyone else keeps theirs.
It’s a stable resting point. Not necessarily the best outcome for everyone — just a point where nobody has a reason to move.
Nash proved something remarkable: every finite game (one with a limited number of players and choices) has at least one Nash Equilibrium. This result is why his name is on it.
Some games have no pure strategy Nash Equilibrium. Think about penalty kicks in football. If the striker always shoots left, the goalkeeper always dives left — and the striker adjusts. You end up with both players randomising to stay unpredictable.
Economists call this a mixed strategy equilibrium — where you randomly choose between options according to specific probabilities. Chiappori, Levitt, and Groseclose (2002) analysed 459 real penalty kicks in professional football and found that the actual behaviour of players matched the mixed strategy Nash Equilibrium remarkably closely. Real people playing a high-stakes game — without knowing any game theory — arrived at the theoretically optimal randomisation.
In a one-shot game, the Prisoner’s Dilemma always ends in mutual defection. But what if you play the same game again and again against the same opponent? Now things change.
If I cheat on you today, you can punish me tomorrow. And if you know I’ll punish you, you have a reason not to cheat. This logic allows cooperation to survive in repeated games — even between purely selfish players.
Robert Axelrod (1984) held computer tournaments of repeated Prisoner’s Dilemma games to find the best strategy. The winner? Tit-for-Tat — cooperate first, then copy whatever your opponent did last round. It’s simple, forgiving, and punishes cheating immediately. It beat every complex strategy in the tournament.
OPEC is a group of oil-producing countries that try to agree on how much oil each member produces. Less production = higher prices = more money for everyone. Sounds good, right?
The problem is a classic Prisoner’s Dilemma:
Between 2014 and 2016, this is basically what happened. Saudi Arabia — frustrated with other members cheating — stopped restricting its own output. Oil prices crashed from $110 per barrel to $27. Every member lost hugely.
So how does OPEC sometimes succeed? Saudi Arabia acts as what economists call a “residual supplier” — it threatens to flood the market if others cheat. That threat can make the long-run cost of cheating high enough that cooperation is rational, especially in repeated interactions where the relationship matters.
This matches the folk theorem perfectly: in a repeated game, cooperation is sustainable when players value the future highly enough and punishment for defection is swift and severe.
Sources: Alhajji, A.F. & Huettner, D. (2000). Energy Journal, 21(3). Axelrod, R. (1984). The Evolution of Cooperation. Basic Books.
Not all games are played simultaneously. In many real situations, one player moves first and the other responds. Think about a new firm deciding whether to enter a market. The existing firm then decides whether to fight (start a price war) or just accept the competition.
We solve these using backward induction — start at the end and work backwards. What would the existing firm do if entry happened? If fighting a price war hurts the incumbent too, the rational choice is to accept competition. Knowing this, the new firm enters. The threat of a price war isn’t credible — so it doesn’t stop entry.
This is why companies sign long-term contracts, build excess capacity, or commit publicly to aggressive pricing — they’re trying to make otherwise empty threats credible.
Two firms (A and B) are each deciding whether to advertise. The payoff matrix (profits, £m) is:
| B: Advertise | B: Don’t | |
|---|---|---|
| A: Advertise | 4, 3 | 6, 1 |
| A: Don’t | 2, 5 | 3, 4 |
Find the Nash Equilibrium. Does either firm have a dominant strategy?
Firm A’s best responses: If B advertises → A prefers Advertise (4 > 2). If B doesn’t → A prefers Advertise (6 > 3). ✅ Advertise is a dominant strategy for A.
Firm B’s best responses: If A advertises → B prefers Advertise (3 > 1). If A doesn’t → B prefers Advertise (5 > 4). ✅ Advertise is a dominant strategy for B.
Nash Equilibrium: (Advertise, Advertise) → payoffs (4, 3).
Notice: if neither advertised, payoffs would be (3, 4) — B would prefer that! This is another Prisoner’s Dilemma structure — both end up advertising and competing away some profit, even though mutual restraint would benefit B.
Explain why the Nash Equilibrium in a one-shot Prisoner’s Dilemma is Pareto inefficient. How can repeated interaction change this outcome?
Why it’s Pareto inefficient: In the Nash Equilibrium both players defect — because defection is a dominant strategy. But both would be better off if both cooperated. So the Nash outcome is not Pareto efficient — we can make both players better off by changing from (Defect, Defect) to (Cooperate, Cooperate). Individual rationality produces a collectively bad outcome.
How repetition helps: When the same game is played repeatedly, players can use past behaviour to reward or punish each other. A Tit-for-Tat strategy — cooperate first, then copy your opponent — creates an incentive to cooperate today because defecting today triggers punishment tomorrow. The folk theorem tells us that as long as players value future payoffs enough (high discount factor), cooperation can be sustained indefinitely as a Nash Equilibrium of the repeated game. This is why ongoing trade relationships, OPEC meetings, and international environmental agreements can sometimes achieve cooperation where one-shot negotiations can’t.
The exponents in Q = AL^α K^β are not just parameters — they are output elasticities, they sum to the returns to scale, and under competition they equal income shares. A complete guide to the Cobb-Douglas production function: marginal products, diminishing returns vs returns to scale, cost minimisation, Solow’s growth accounting, and the declining labour share — with worked examples and exam technique.
Statistics · EstimationConfidence Intervals ExplainedPoint estimates give you one number — confidence intervals tell you how much to trust it. Here's exactly what they mean and how to calculate them.You poll 500 voters and find 52% support a particular policy. But you...
Every time you would have happily paid more than you did, you pocketed consumer surplus. A complete guide to measuring welfare with integration: consumer and producer surplus, total surplus and market efficiency, deadweight loss, why tax loss grows with the square of the rate, and non-linear demand curves — with worked examples and exam technique.
The exponents in Q = AL^α K^β are not just parameters — they are output elasticities, they sum to the returns to scale, and under competition they equal income shares. A complete guide to the Cobb-Douglas production function: marginal products, diminishing returns vs returns to scale, cost minimisation, Solow’s growth accounting, and the declining labour share — with worked examples and exam technique.
Statistics · EstimationConfidence Intervals ExplainedPoint estimates give you one number — confidence intervals tell you how much to trust it. Here's exactly what they mean and how to calculate them.You poll 500 voters and find 52% support a particular policy. But you...
Every time you would have happily paid more than you did, you pocketed consumer surplus. A complete guide to measuring welfare with integration: consumer and producer surplus, total surplus and market efficiency, deadweight loss, why tax loss grows with the square of the rate, and non-linear demand curves — with worked examples and exam technique.