Go and look

Almost everything we argue about has a number in it; almost every number we have is a proxy


Shut up and measure.

act65

Most arguments I get into have, somewhere inside them, a claim that could be checked — that harsher sentences deter, that the benefit makes people work less, that the treatment isn’t worth what it costs. The argument runs for hours anyway, and at no point does anyone ask what number would settle it, who would have to go and collect it, or what we would all do differently if it came back the other way. 1

One: if it’s an empirical question, stop arguing and go and look

The history of philosophy contains a long run of confident conjectures about life and mind that were not settled by argument. They were settled, often centuries later, by somebody measuring.

Heavier bodies fall faster: argued from first principles, believed for the better part of two thousand years, and undone by dropping two lead balls off a church tower in Delft and listening for one sound instead of two. 2 Life arising spontaneously from decaying matter: an ancient and reasonable-sounding position, dismantled by Redi covering some meat with gauze and, much later, Pasteur’s swan-necked flasks. Living things possessing a vital force that no ordinary chemistry could produce: a serious philosophical position, worn down by Wöhler’s urea and then by Buchner getting fermentation to run in a cell-free extract, with no living cell present to supply the vitality. 3 The seat of thought: Aristotle said the heart, and people argued about it for two millennia until Broca did an autopsy on a man who had lost his speech and found the lesion. 4

The modern version of this is intelligence, and we are still doing it. Every so often someone says a task requires genuine understanding — chess, then Go, then translation, then whatever is next — and this is stated as a philosophical claim about the nature of mind when it is actually a claim about what a machine will score on a task. It gets settled the same way: someone builds the thing and measures.

The pattern is not that the philosophers were stupid. Several of these positions were the reasonable inference from what was known. The pattern is that the question was empirical the whole time, and the arguing part contributed roughly nothing to answering it.

My exemplar for the opposite temperament is Alexander von Humboldt, and he is an exemplar partly because he took it much too far. He crossed South America between 1799 and 1804 with something like forty instruments and measured essentially everything he walked past: air temperature, water temperature, barometric pressure, the magnetic field, his own pulse, the blueness of the sky — there is an instrument for that, a cyanometer, and he carried one. Told about the electric eels in the pools near Calabozo, he had horses driven into the water to draw off their charge and then waded in and picked the eels up, recording the shocks, the pain through his knees and joints, and the fact that he was no use to anyone for the rest of the day. He had already spent a couple of years applying electrodes to cuts and blisters raised on his own back, to see what galvanism did to nerve and muscle. 5

That is a temperament rather than a method, and what it produced was structure nobody had been arguing about. Out of the readings taken up the slopes of Chimborazo came a picture of vegetation banded by climate rather than geography, and from there the idea of nature as a single interconnected system. At Lake Valencia he recorded a falling water level and blamed the plantations clearing the forest around it. I don’t think you argue your way to either of those. You carry the barometer up the mountain, and occasionally you get in the water with the eels.

The same thing happens with policy, at a much faster tempo and with people’s lives inside it. In the 1990s the received wisdom on multi-drug-resistant tuberculosis in poor countries was that treating it was not cost-effective: somewhere around two to three thousand dollars a patient, against ten or twenty for ordinary TB, for a two-year regimen that somebody would have to supervise daily and that most patients were expected to abandon. So the money should go elsewhere.

That is a ratio, and both halves of it were assumptions. Paul Farmer and his colleagues at Partners In Health went at the bottom half first, treating the patients in Carabayllo and counting the cures. What made the cures possible was not a clinical advance but a staffing one: community health workers — in Haiti they are called accompagnateurs, from accompaniment — mostly local people rather than clinicians, going to each patient’s house every day to watch them take the pills, for as long as the regimen ran. Reported cure rates came out around 80%, in a slum, in a population everyone had written off. 6

The top half moved afterwards, and moved further. Most of the second-line drugs were off patent and cheap to make; they were expensive because almost nobody bought them, so nobody made them at scale — a market failure rather than a manufacturing cost, which is how Farmer’s group wrote it up. Consolidating the demand and getting more suppliers qualified cut regimen prices by something between half and almost everything, depending on the drug: cycloserine went from $3.38 a capsule to 14 cents. 7

The received wisdom had not made an arithmetic error. It had taken both numbers in the ratio as given, when both were consequences of choices nobody had revisited, and the unrevisited numbers decided who got treated.

So: where a claim is testable, debating it looks like close to the worst available use of the time, and going to look like close to the best. That is the part I’m confident about. It is also, in practice, almost never the situation I’m in.

Two: the data is never all in, and a bad number is worse than none

Real decisions get made before the measurement exists, with a proxy standing in for the thing we actually care about. And a proxy has something that an honest argument doesn’t: the authority of a number, with none of the caveats attached.

The clean example is QALYs. Somebody has to decide which treatments a health system funds, that decision is a rationing decision whether or not anyone says so out loud, and quality-adjusted life years at least make it explicit and comparable. That is a real gain. But the quality weights are a moral judgement wearing accountant’s clothes: a year of life for a disabled person is worth less than a year for an abled one by construction, because that is what the weights say. The arithmetic is then impeccable. 8 The counting is not what’s wrong: choosing what to count was the whole ethical decision, and it got made in a footnote.

The uglier example is craniometry. Through the nineteenth century a lot of serious people filled skulls with mustard seed, and later with lead shot for consistency, poured the contents into a graduated cylinder, and published the volumes grouped by race — and, elsewhere, brain weights grouped by sex. Europeans came out on top, Africans at the bottom, women below men. The instruments were real and the protocols were written down. What did the actual work was the step that never got measured at all: that skull volume stands in for intelligence, and intelligence for human worth. Lombroso went further and read criminality off the bumps. The ranking that came out was, in each case, the ranking the measurer walked in with.

It is worth adding that the best-known modern debunking of this, Gould’s, was itself challenged by people who went back and re-measured the skulls, and that challenge has been disputed in turn. 9 Even the correction is contested. A number is not evidence that anyone was careful.

And the failure mode that names itself: measure what is measurable, then treat what you couldn’t measure as unimportant, then treat it as nonexistent. Vietnam-era body counts are the standard illustration. 10 What can’t be counted doesn’t merely get left out of the model; it gets left out of the argument.

The rest are variations, and duller: the proxy measures the wrong dimension and quietly changes who gets helped; the summary statistic turns out not to be one thing; the outcome you measured cleanly was mostly luck; the number exists, so somebody optimises it.

Is a bad metric really worse than none? I think often yes, and the reason is that it ends the argument instead of informing it. Without the number, people argue, and they know they’re arguing. With it, the disagreement looks settled, the value judgement has been laundered into arithmetic, and the burden of proof has quietly moved onto whoever objects.

Where that leaves me

Not with a method. Something more like: the first question is always whether this is the kind of thing that could be settled by going and looking, because if it is, then the argument I’m having is a waste of a good afternoon. And if it isn’t — if the best available number is a proxy standing in for the thing — then the number is a contribution to the argument rather than the end of it, and the choice of what to count is the part to argue about.

And there is a more mercenary reason to want the number anyway: once something is counted, you can price it, regulate it, or hold somebody to it. Nobody fixes the broken paving because the cost of my ankle appears on nobody’s books.

Both kinds, on this blog

What Makes a Good Politician? August 1, 2026
The one standard we can all agree on, and how our MPs score against it

Ballot Structure from the Preference-Dependency Graph July 17, 2026
Which questions should a ballot ask together? The votes can tell us.

Calculating the Mean October 10, 2025
I thought means were simple. I was wrong.

Beyond Earning to Give October 10, 2025
The EA Road to AI

Ikram and his tricks October 8, 2025
A Tale of Clay and Calculation

Our Attention is a Multi-Billion Dollar Asset September 21, 2025
Who's Cashing In?

Committed to Fidelity July 3, 2025
A Unified Model for Multi-Fidelity Bandits

Freedom-Adjusted Life Years July 2, 2025
Are We Solving the Right Problems?

Affirmative Action's Simplistic Approach May 20, 2025
Using a 1-Dimensional Solution for a Multi-Dimensional Problem

Luck vs skill in poker and life December 20, 2024
Why does poker have many rounds?

A tripping point May 21, 2022
I tripped over and have something to say about it.

The future of environmental sciences? February 7, 2019
What could the profession of environmental engineering look like?

Wish list of psychological metrics November 5, 2018
What if we could accurately measure {INSERT}?

What are these power-compare websites for? March 26, 2018
An example of how to regulate markets to remove inefficiencies.

Policy search engine December 10, 2017
Motivating a tool for tracking the effects of policies.

A guide to regulation October 31, 2017
If we measure things, we can regulate them, fairly.


  1. With apologies to David Mermin, who wrote “shut up and calculate” about people arguing over what quantum mechanics means, and to whom the sentiment at the top of this page owes rather a lot. 

  2. Simon Stevin and Jan Cornets de Groot, Delft, around 1586, reported in Stevin’s De Beghinselen der Weeghconst (1586). Galileo’s Leaning Tower version is probably a later story; his inclined-plane experiments are the real ones. 

  3. Friedrich Wöhler synthesised urea from ammonium cyanate in 1828, and Eduard Buchner produced fermentation in a cell-free yeast extract in 1897. Historians of chemistry are clear that Wöhler’s synthesis did not kill vitalism on its own — the potted version overstates it — but the direction of travel was measurement, not argument. 

  4. Paul Broca, 1861, on the patient known as Tan. Alcmaeon and the Hippocratic writers had picked the brain long before, so this was less a discovery than the end of a long stalemate that argument had not been able to break. 

  5. I got most of this from Daniel Kehlmann’s Measuring the World (2005), which is a novel and takes liberties with both Humboldt and Gauss; Andrea Wulf’s The Invention of Nature (2015) is the non-fiction version. The eels are in Humboldt’s own account of Calabozo, March 1800, and the self-experiments in Versuche über die gereizte Muskel- und Nervenfaser (1797). Humboldt’s Essay on the Geography of Plants (1807) has the Chimborazo diagram. 

  6. The Peruvian programme ran as Socios En Salud; the cohort was written up by Mitnick et al., “Community-Based Therapy for Multidrug-Resistant Tuberculosis in Lima, Peru”, New England Journal of Medicine 348 (2003). Tracy Kidder’s Mountains Beyond Mountains (2003) tells the story from the inside. Accompaniment is Farmer’s word for the approach rather than just the job title, and his argument for it is not only that it works — it is also what you would want if you were the patient. 

  7. Gupta, Kim and Farmer, “Responding to Market Failures in Tuberculosis Control”, Science 293 (2001), report reductions of 48–97% across a regimen; the WHO announcement of the Green Light Committee prices the same July put the top figure at 94%, and the cycloserine numbers are from that round of negotiations. The mechanism was pooled demand and more qualified suppliers, not a discovery that the drugs had been cheap all along — and it followed the cure-rate result rather than preceding it. 

  8. This is a live enough objection that US law restricts the use of cost-per-QALY thresholds in Medicare coverage decisions, on disability-discrimination grounds. 

  9. Stephen Jay Gould, The Mismeasure of Man (1981), argued that Samuel Morton’s skull measurements were unconsciously fudged; a 2011 re-measurement argued Gould was the one doing the fudging; later work has disputed that. I have not read enough to have a view on who is right, which is rather the point. 

  10. Usually called the McNamara fallacy, after the Vietnam-era body counts. The formulation — measure what can be measured, disregard what can’t, presume what can’t be measured isn’t important, then that it doesn’t exist — is usually credited to Charles Wyllys Adams and was popularised by Daniel Yankelovich in 1972.