Global coordination problems via game theory

Why AGI is not like climate change

Revised 11 Jun 2026, 7 Aug 2026


Some global coordination problems get easier with time, and the reason is in the payoffs. Overfishing and climate change look like Prisoner’s Dilemmas in the short run. Extend the horizon and the numbers move: a healthy fishery is worth far more than one good season, so mutual cooperation stops being the second-best outcome and becomes the best one. The Prisoner’s Dilemma becomes a Stag Hunt. That transformation is real, and it’s what the rest of this post is built on.

What it doesn’t do is solve anything by itself. A Stag Hunt still has two equilibria, and mutual defection is one of them. Getting a whole world to the good one is the hard part of governance — and it often doesn’t happen in time.

AGI breaks the pattern. The conditions that drive the transformation elsewhere — depleting resources, repeated interactions, visible consequences — don’t apply, or apply too slowly. What’s left is a Prisoner’s Dilemma that doesn’t dissolve on its own, with a defection prize unlike anything we’ve faced — and it stays one for exactly as long as the people racing believe they can race and win.

Every outcome below is scored 1 to 4 for each party — 4 is that party’s best outcome, 1 its worst — and written as (Player A, Player B, Humanity). Humanity is a scorecard, not a player: it has no strategies, never moves, and its number just records how good the outcome is for everyone not at the table.

Two orderings do all the work.

In a Prisoner’s Dilemma, each player ranks the four outcomes this way: best is defecting while the other cooperates (4); next, both cooperate (3); then both defect (2); worst is cooperating while the other defects (1). Whatever the other player does, you score higher by defecting. So both defect and both land on 2, when both could have had 3.

A Stag Hunt swaps the top two: both cooperating (4) now beats defecting on a cooperator (3). Defection stops being the automatic move, and the game has two equilibria instead of one.

That’s enough to check my work as we go, which you should.

The Standard Pattern: From Short-Term Dilemma to Long-Term Cooperation

The core of most global coordination challenges lies in overcoming short-term incentives to secure a better long-term future. The goal of governance is to build the trust needed to make this shift.

Example 1: Overfishing

In the short term, the game is a Prisoner’s Dilemma. The immediate incentive is to overfish for a quick profit, especially if you fear others will do the same. This logic pushes everyone toward the tragedy of the commons (Hardin, 1968).

Cells read (You, Others, Humanity). Collapsing every other fisher into a single “Others” player is a real simplification — it removes exactly the free-riding structure that makes commons problems hard in practice — but it keeps the matrix legible.

  Others: Fish Sustainably Others: Overfish
You: Fish Sustainably (3, 3, 3) (1, 4, 2)
You: Overfish (4, 1, 2) (2, 2, 1) - Mutual Ruin

Read two cells to get the hang of it. Bottom right, (2, 2, 1): everyone overfishes, so you and the others score 2 — third of four — and humanity scores 1, its worst. Top right, (1, 4, 2): you fished sustainably while the others overfished, so you take your worst outcome and they take their best. Whichever column you look down, overfishing scores higher for you — so you overfish, and so does everyone else, landing all of you on 2 when mutual restraint would have paid 3.

Humanity gets a 3, not a 4, at mutual cooperation: sustainable fishing under short-run pressure is good, but it’s not as good as the standing healthy fishery in the next matrix. That’s the fixed cross-matrix scale doing its job.

In the long term, the game becomes a Stag Hunt. A healthy, sustainable fishery is vastly more valuable than any short-term gain. The (4, 4, 4) outcome of mutual cooperation becomes the best prize for the players and for Humanity.

  Others: Fish Sustainably Others: Overfish
You: Fish Sustainably (4, 4, 4) - Sustainable Bounty (1, 3, 2)
You: Overfish (3, 1, 2) (2, 2, 1)

The mechanism is concrete: repeated interactions let cooperators punish defectors, and the resource degrades as defection accumulates, shrinking the defector’s prize — the boat that overfishes a depleted stock lands much less than one that overfished a healthy stock, which is why temptation drops from (4) to (3). Note the lone conservationist still scores (1), below mutual collapse. Unilateral restraint doesn’t save the stock: you lose twice, because the fishery goes anyway and somebody else landed the last of it.

Mutual overfishing (2, 2, 1) remains an equilibrium of the long-term game, as promised. What changes is that cooperation becomes self-enforcing once reached. The job of fisheries institutions — treaties, quotas, enforcement — is equilibrium selection, not payoff engineering: making (4, 4, 4) the focal point, so everyone believes everyone else is aiming at the same cell (Schelling, 1960).

Example 2: Climate Change

Climate has the same structure, so I won’t redraw it. In the short run the matrix is identical to the overfishing one above, with “pollute” in place of “overfish” and nations in place of boats: abatement is costly, free-riding pays, and everyone lands on (2, 2, 1). In the long run it’s the same Stag Hunt — the catastrophic costs of a runaway climate are supposed to make a stable planet the ultimate prize, with (4, 4, 4) as the cell everyone would like to coordinate on. Same mechanism, same two equilibria, same problem of getting to the better one.

But does anyone actually take the long view?

Rarely. That’s worth separating from the claim above, because they’re different claims. The transformation is a fact about the payoffs: extend the horizon and mutual cooperation becomes the best outcome available. Whether anyone acts on the extended horizon is a separate question, and there the record is bad.

Northern cod collapsed in 1992 and hasn’t recovered thirty years on. The share of assessed marine stocks fished at biologically unsustainable levels has risen from around 10% in the mid-1970s to roughly 37% in the FAO’s latest assessment (FAO, 2026). Emissions have risen through three decades of climate diplomacy: Kyoto was never ratified by the US and Canada withdrew, Paris is non-binding.

So the Stag Hunt is available and we keep landing on its bad equilibrium. That’s an equilibrium-selection failure, not a payoff failure — and selection is what Elinor Ostrom’s work is about. She documented hundreds of common-pool resources managed successfully for centuries; the ingredients are clear boundaries, monitoring, graduated sanctions and recognised rights to organise, applied before collapse rather than after (Ostrom, 1990). Her own view was that these don’t scale to global commons, which is why she argued for polycentric climate governance over a grand treaty (Ostrom, 2010). Barrett cuts the same way: environmental agreements tend to self-enforce only when the gains from cooperation are small (Barrett, 2003). Montreal is the one unambiguous success and it fits him — cheap abatement, available substitutes, few producers.

None of which rescues my optimism about fisheries or climate. But notice what kind of failure it is. In both cases the good equilibrium exists, we are simply bad at reaching it, and Ostrom tells us roughly what would help. The claim in the next section is worse than that: for AGI there is no horizon at which the payoffs turn cooperative, so there is no good equilibrium for monitoring and sanctions to select. Missing a target is a different problem from not having one.

Why AGI Is Different: A Permanent Dilemma

The transformations that rescue other coordination problems don’t have an obvious analogue for AGI:

  • No resource depletion. Overfishing destroys the fishery, shrinking the defector’s prize. AGI capability doesn’t deplete with use — if anything, a first-mover compounds their lead.
  • No repeated rounds. Climate cooperation accrues over decades of small, reversible decisions. The most pessimistic AGI race models look more like a single round: whoever gets there first locks in (Armstrong et al., 2016; Naudé & Dimitri, 2020).
  • No visible degradation. A collapsing fishery is something everyone can see. A capabilities lead, an alignment failure mode, or an internal safety culture is largely invisible to the other player.

The strongest objection is to that second bullet. The AGI decision isn’t isolated: it sits inside a dense repeated relationship — trade, chips, tariffs, standards bodies, talent mobility, publication norms, joint safety evaluations — and issue linkage can sustain cooperation on a question that is one-shot in itself (Askell et al., 2019). I think that’s insufficient rather than wrong. Linkage works when the linked stakes are comparable to the stake in question, and nothing in the trade relationship is worth what the winner of this one believes they are getting.

Strip the rest away and the whole thing rests on one belief: that getting there first is better than cooperating. Not that it’s true — that the people making the decision believe it. Hold that belief and everything below follows. Drop it and none of it does.

So the same payoffs stay on the table at long horizons:

The AGI Payoff Matrix (Player A, Player B, Humanity)

  Player B: Cooperate Player B: Defect
Player A: Cooperate (3, 3, 4) - Mutual Safety (1, 4, 2) - Sucker & Temptation
Player A: Defect (4, 1, 2) - Temptation & Sucker (2, 2, 1) - Mutual Ruin

Note what’s happened here: the player payoffs in this matrix are identical to the short-run overfishing matrix above. That’s the point — and also the limit of what a payoff matrix can show. The matrix doesn’t prove AGI is different; it just records that I think the same short-run game persists at long horizons, where in fisheries it doesn’t. The argument for that is in the three bullets above, not in the numbers.

Or is it Chicken?

This matrix encodes one strong assumption: that being beaten to AGI (1) is worse for a player than a mutual race (2). That is only true if players’ payoffs are purely positional — if they value relative standing and don’t internalise their own share of the catastrophe. Nations and firms often do behave this way. But the whole force of the argument below is that a botched race is existentially catastrophic, and racers are in the set of people who die. If the actors take that risk personally, (1) and (2) swap, the ordering becomes Temptation > Reward > Sucker > Punishment, and the game is Chicken:

  Player B: Cooperate Player B: Defect
Player A: Cooperate (3, 3, 4) - Mutual Safety (2, 4, 2)
Player A: Defect (4, 2, 2) (1, 1, 1) - Mutual Ruin

Neither player has a dominant strategy now: race if they cooperate, cooperate if they race. There are two asymmetric pure equilibria — one racer, one abstainer — plus a mixed equilibrium that puts positive probability on the catastrophe. The risk becomes endogenous to the game rather than an assumption about it.

Chicken is in some ways the worse game. It rewards visible, irrevocable commitment — the player who credibly rips out their own steering wheel wins (Schelling, 1960) — which inverts the prescription: “build common knowledge” becomes the wrong move, because what you broadcast is exactly what a committed racer wants you to have. Which game we’re in depends on whether decision-makers price their own deaths into the payoff. I’ll take the positional version for the rest of the post, because I think it describes how these decisions are actually made. But it’s an assumption, and it’s carrying the argument.

The misalignment is an assumption, and it’s the important one

Compare mutual cooperation across the games:

  • Long-term overfishing and climate: (4, 4, 4) — players and humanity all get their best available outcome simultaneously.
  • AGI: (3, 3, 4) — humanity’s best, but each player still sees a higher (4) they could have grabbed by defecting unilaterally.

That gap is the misalignment, and it isn’t derived — it’s assumed, in the decision to give players payoffs that don’t include humanity’s. Defensible, since firms answer to shareholders and states to citizens and neither is “humanity”. But it’s the input, not the output. Nor is the gap distinctive to AGI: a player seeing a better payoff they could grab is just the definition of a Prisoner’s Dilemma. What’s distinctive is the claim that this one never transforms.

This is the deeper version of the AI alignment problem. Aligning the model to its operator isn’t enough if the operators themselves aren’t aligned to humanity. See here for more on this argument.

So defection is dominant for both, and unlike fisheries the (4) doesn’t shrink as the game plays out. Repetition could still sustain cooperation in a Prisoner’s Dilemma whose payoffs never change — the folk theorem — but only if players expect enough future rounds to matter, which the second bullet above denies. Prize-shrinkage and repetition are two escape routes; I’ve closed one.

The escape hatch: the prize may not be real

Up to here the payoffs have been ordinal — ranks, not magnitudes. To talk about expected payoffs we need cardinal utilities, so for this section read the numbers as rough utilities rather than ranks.

The strongest counterargument to this whole frame is that the unilateral (4) is conditional on solving alignment, and racing makes solving alignment less likely. That is small enough to put numbers on. Let $p$ be the probability that a racer loses control of what it builds, $W$ the value of winning with a controlled AGI, $C$ the value of mutual cooperation, and $L$ the value of the catastrophe. Racing while the other cooperates beats cooperating iff

\[(1-p)W + pL > C \iff p < \frac{W - C}{W - L}\]

That right-hand side reads more easily than it looks. The numerator is what you gain by winning rather than cooperating; the denominator is the whole span from your best outcome to your worst. So the threshold is the fraction of the total stakes that winning-instead-of-cooperating actually buys you. Believe winning is vastly better than any negotiated outcome and that fraction is large, so almost any risk is worth running. Believe cooperation gets you most of what winning would, and it’s small — a slim chance of catastrophe is enough to make racing a bad bet.

Rough numbers make it concrete. Say winning with a controlled AGI is worth 100 to you, a cooperative settlement 70, and losing control $-1000$. The threshold comes out at 30/1100, under 3%: believe there’s more than a 3% chance you lose control and racing stops paying. Make cooperation less attractive — say it leaves you sidelined at 10 — and the threshold only rises to about 8%. It stays small either way, because the catastrophe term dominates everything else in the expression.

Which is where this reconnects to the Chicken question. That arithmetic only works if you price the catastrophe into your own payoff. A purely positional player — one who cares about relative standing and treats losing control as just another way of not winning — faces a far higher threshold, and races almost regardless of $p$.

So why do decision-makers put $p$ below the threshold? Because of what the prize is believed to deliver:

  • Existential security. The first actor to build a controllable superintelligence could end strategic competition globally and permanently.
  • Economic singularity. Whoever controls the first AGI captures most or all of the value created by automation.
  • Fear of irrelevance. The (1) is not just a strategic setback; it’s the risk of your nation, culture, or company being permanently sidelined.

All three inflate $W$ and deflate $C$, which pushes the threshold up and makes racing look rational at higher risk. All three are contested in detail — and, as above, they don’t have to be true, only believed. Which is a reason to be careful with the framing: Cave and ÓhÉigeartaigh argue the race narrative is partly self-fulfilling, that writing posts like this one helps create the game they describe (Cave & ÓhÉigeartaigh, 2018). I don’t have a clean answer beyond trying to be precise about which parts are assumed and which are derived.

So the post reduces to one inequality, and everything under “what could change the game” below is an intervention on $p$, on beliefs about $p$, or on $W$. Han and colleagues derive essentially this threshold properly, as a function of the risk-to-speed ratio, and show that the AI race is a dilemma on one side of it and a coordination game on the other (Han et al., 2020).

Cross the threshold and the matrix doesn’t merely soften; it changes game. The temptation payoff falls below the reward payoff, defection stops being dominant, and mutual cooperation becomes the payoff-dominant equilibrium of a two-equilibrium coordination game — a Stag Hunt. Cooperation becomes achievable, not automatic. Mutual defection is still sitting there as an equilibrium of the new game.

This is the belief update AI safety research is trying to force into common knowledge. Unlike fisheries, the transformation here isn’t automatic — it has to come from evidence and persuasion rather than from the resource depleting on its own.

What about nuclear weapons?

The obvious objection: nuclear weapons share most of these features. A permanent prize, existential stakes, a first-mover advantage, no Stag Hunt transformation — and yet we got partial cooperation through MAD, the NPT, and test ban treaties. What that section of history is really about is the security dilemma: measures each side takes for its own safety are indistinguishable from preparations to attack, so both end up less safe (Jervis, 1978).

So why is AGI different from nukes? The honest answer is that AGI is nukes with several of the stabilizers missing:

  • No second strike. MAD works because retaliation is credible — you can absorb a first strike and respond. If Player A reaches aligned superintelligence first, Player B has no analogous retaliation capacity.
  • Hard to verify. Nuclear tests produce seismic, atmospheric, and supply-chain signatures. AGI development happens in datacenters and ships as model weights; verification regimes are still largely speculative.
  • Faster timelines. Decades of nuclear arms racing left room to build diplomatic infrastructure around it. AGI timelines may compress this to years.
  • Self-destruction, not deterrence. The MAD-equivalent for AGI isn’t “I retaliate if you launch” but “your own creation destroys you” — which deters less because it’s probabilistic and contested.

AGI isn’t unique in being a Prisoner’s Dilemma over an existential prize. It’s that the institutional and physical stabilizers that turned nukes into a (precarious) equilibrium aren’t here yet.

What could change the game

If the transformation doesn’t happen on its own, what could induce it? (There is a whole research agenda on this question (Dafoe, 2018).)

  • Verifiable compute monitoring. The closest historical analogue is nuclear test verification — making racing detectable changes the payoff for unilateral defection. Compute is one of the few legible inputs to AGI development.
  • Shared evidence of alignment failure. Concrete demonstrations that powerful models fail in ways their operators didn’t intend collapse the belief that the racer cleanly captures (4).
  • Actually measuring $p$. The probability a racer loses control is currently guessed, not estimated. Evaluations, incident reporting and a public record of how often frontier systems behave in ways their developers didn’t intend would make it a number. Given how low the threshold sits, this may be the cheapest intervention available — if the true value is above it, measuring well changes the game with no treaty required. The caveat is symmetric: if it’s below, better measurement makes racing look more attractive. It’s the one intervention that can backfire by working.
  • Mutual vulnerability. Cyber, model exfiltration, and open-weights diffusion mean a “winner” is unlikely to stay a winner. Making this legible shifts the perceived durability of the prize.
  • Common-knowledge constraints. Treaty-style commitments on training compute or specific capabilities — even partial — change the equilibrium, because each player needs to know that the other knows.

One caveat cuts against the first of those. In Armstrong, Bostrom and Shulman’s race model, risk rises with the number of teams and with information: the more each team knows about the others’ capabilities and its own, the greater the danger, because a team that knows it is behind takes bigger risks (Armstrong et al., 2016). Monitoring that reveals relative standings is not automatically safe. Monitoring that verifies compliance with a ceiling without revealing who is ahead is a different, and harder, design problem.

None of these guarantees a transformation. Each chips at one of the assumptions holding the Prisoner’s Dilemma in place.

The shape of the problem

Most global coordination challenges are supposed to resolve when actors come to see a larger mutual prize and shift from Prisoner’s Dilemma to Stag Hunt. On the evidence, they often don’t. AGI is worse placed than most: it has neither the resource dynamics nor the repeated interactions that are meant to drive that shift, and unlike nuclear weapons the stabilizers that produced even precarious cooperation aren’t here.

But the pessimism is conditional, and that’s the useful part. The misalignment between players and humanity is an assumption about whose welfare is in the payoffs, not a discovery. The Prisoner’s Dilemma holds only while decision-makers put the probability of losing control below $(W-C)/(W-L)$. And if they price the catastrophe into their own payoffs rather than treating it as someone else’s problem, the game isn’t a Prisoner’s Dilemma at all — it’s Chicken, which is not better. What holds all three of those in place is belief, which is why the fight is over what becomes common knowledge among the people racing.

What this model assumes

A model this small buys clarity by lying about the details. The load-bearing lies:

  • Two players. There are several states and a dozen-plus frontier labs. Cooperation needs everyone, defection needs one — the unilateralist’s curse (Bostrom et al., 2016). Two players understates the problem.
  • States are the players, labs are their instruments. They aren’t interchangeable, and only the state reading makes the “existential security” argument work.
  • A 2×2 matrix is the wrong object. The real choice is a continuous safety-versus-speed dial, made sequentially against a stochastic finish line, under poor information about the other side — which is what the race literature actually models. Properly this is a Bayesian game about beliefs.
  • A single decisive moment. Whoever gets there first locks it in. If lock-in doesn’t hold, the temptation payoff isn’t really a (4).

Bibliography

  1. Hardin, G. (1968). The Tragedy of the Commons. Science, 162(3859), 1243–1248. https://doi.org/10.1126/science.162.3859.1243
  2. Schelling, T. C. (1960). The Strategy of Conflict. Harvard University Press.
  3. FAO. (2026). The State of World Fisheries and Aquaculture 2026. Food and Agriculture Organization of the United Nations.
  4. Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press.
  5. Ostrom, E. (2010). Polycentric systems for coping with collective action and global environmental change. Global Environmental Change, 20, 550–557. https://doi.org/10.1016/j.gloenvcha.2010.07.004
  6. Barrett, S. (2003). Environment and Statecraft: The Strategy of Environmental Treaty-Making. Oxford University Press.
  7. Armstrong, S., Bostrom, N., & Shulman, C. (2016). Racing to the precipice: a model of artificial intelligence development. AI & Society, 31(2), 201–206. https://doi.org/10.1007/s00146-015-0590-y
  8. Naudé, W., & Dimitri, N. (2020). The race for an artificial general intelligence: implications for public policy. AI & Society, 35, 367–379. https://doi.org/10.1007/s00146-019-00887-x
  9. Askell, A., Brundage, M., & Hadfield, G. (2019). The Role of Cooperation in Responsible AI Development. ArXiv Preprint ArXiv:1907.04534. https://arxiv.org/abs/1907.04534
  10. Cave, S., & ÓhÉigeartaigh, S. S. (2018). An AI Race for Strategic Advantage: Rhetoric and Risks. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, 36–40. https://doi.org/10.1145/3278721.3278780
  11. Han, T. A., Pereira, L. M., Santos, F. C., & Lenaerts, T. (2020). To Regulate or Not: A Social Dynamics Analysis of an Idealised AI Race. Journal of Artificial Intelligence Research, 69, 881–921. https://doi.org/10.1613/jair.1.12225
  12. Jervis, R. (1978). Cooperation Under the Security Dilemma. World Politics, 30(2), 167–214. https://doi.org/10.2307/2009958
  13. Dafoe, A. (2018). AI Governance: A Research Agenda. Future of Humanity Institute, University of Oxford.
  14. Bostrom, N., Douglas, T., & Sandberg, A. (2016). The Unilateralist’s Curse and the Case for a Principle of Conformity. Social Epistemology, 30(4), 350–371. https://doi.org/10.1080/02691728.2015.1108373