What Makes a Good Politician?
The one standard we can all agree on, and how our MPs score against it
August 1, 2026
Here is a New Zealand Cabinet minister, in the House, on the record, describing another member:
“If this guy opened a funeral parlour, nobody would die. He can screw up anything—” — David Seymour (Hansard, 10 Mar 2026)
And here is another minister, asked in the House to explain what systemic bias is:
“Well, systemic bias is bias that’s systemic.” — Mark Mitchell (Hansard, 29 Aug 2024)
I came across a run of clips like this — first from the UK Parliament, then ours — and the thing that surprised me wasn’t that politics is rude. It was the setting. This is not a campaign rally or a late-night panel show. This is them at work, in the chamber, doing the job we elected them to do, and the record of it is kept forever.
Is that really how they behave while working? I couldn’t tell from clips, because clips are selected to be bad. So I went and measured it.
What should we expect?
Before measuring anything I had to decide what “good” even means here, and the distinction that makes it tractable is between what a politician believes and how they conduct themselves.
What they believe is legitimately contested. Capital gains tax, the road, the tunnel — that’s what elections are for. I can’t score it without just scoring “agrees with me,” and neither can anyone else.
How they conduct themselves is a different kind of question — and if government and legislation really are matters of reason and judgment, as Burke told the voters of Bristol in 1774, then how an MP argues isn’t decoration on top of the job. It is the job. We don’t send them to Wellington to hold opinions; we send them to sit through the select committee, read the officials’ advice, hear the submitters, and reason their way to a decision. Legislate, scrutinise the executive, represent. Everything else is downstream.
The one standard we can all agree on
Here’s the part I keep coming back to.
Take nine plain expectations of conduct. Answer the question you were actually asked. Say something, rather than something-shaped. Get your facts right. Argue from evidence rather than fallacy. Do what you said you’d do. Attack the argument, not the person.
A good politician scores 100 on every one of them.
Not “better than the other lot.” Not “about average for Parliament.” One hundred. Every one is a floor rather than a stretch goal, and I don’t think there’s a voter in the country who’d call any of them unreasonable.
That’s what makes conduct worth measuring when policy isn’t. You and I might disagree completely about tax, immigration, Te Tiriti, the lot — and still agree, without argument, that the person making our case should answer the question and tell the truth while doing it.
Two honest qualifications, since I’m claiming this is the uncontested part. The first is that “answer the question” isn’t uncontested even in the rules: New Zealand’s Standing Orders require a Minister to address the question, not to answer it, and the Speaker rules on that distinction regularly. When I score forthrightness I’m holding MPs to a stricter standard than the House does. The second is that any rubric written down has a house style, and mine has one — it rewards figures, mechanisms and named policies, and marks down hyperbole and exaggeration. That’s a preference for the technocratic voice over the moral or the oratorical, and it isn’t neutral. I’d defend it as the right register for a legislature. I wouldn’t pretend I hadn’t chosen it.
And it’s worth being blunt about why, because this isn’t a preference for good manners. They are representing you. That’s the whole arrangement: they carry your interests into a room you’re not in, and they speak with your authority when they get there. When the MP you voted for sneers at the other side, ducks a question, or reaches for a number nobody checked, that is done in your name.
So the question isn’t the comfortable one — whether you think politicians in general are a bit much. It’s narrower and it’s aimed at you. How do you want your side to win? Is it fine for them to fight dirty, as long as they’re fighting for something you believe in? Is a lie acceptable if it advances your goals?
Almost nobody, asked that directly, says yes. And no party campaigns on dodging or on getting the facts wrong — every one of them claims the opposite. It is the one standard in politics that isn’t up for grabs.
So it should be possible to check. And mostly it isn’t — fact-checks are one-off and forgotten inside a week, and there’s no durable record of how an MP conducts themselves on the boring Tuesdays, only the clips that go viral. That’s the gap I tried to fill.
Measuring it
NZ Politician Scorecards gives every MP a trading card. Nine attributes, scored 0–100, all oriented so higher is better, all measuring conduct rather than policy:
- Forthrightness — do they answer the question actually asked?
- Strength — do promises turn into implemented policy?
- Veracity — do their factual claims hold up?
- Authenticity — do their words in the House match their words outside it?
- Divination — are their predictions specific enough to be checked, and plausible on the evidence at the time?
- Charisma — can they persuade across the aisle?
- Civility — attacking the argument, or the person?
- Rigor — evidence and valid inference, or fallacy?
- Specificity — is there content in the sentence, or is it a platitude?
Scrape public statements, have an LLM pull out the ones that bear on an attribute and score each against a fixed rubric, average up to a card. The part I care about is that every score links back to the statements behind it — click any stat and you get the quotes, each with its own score, the reasoning, the date, and a link to the source. The number is a shrunk average of the evidence you’re looking at: thin evidence gets pulled toward the population middle on purpose, so a single flattering quote can’t make a card. If you think it’s wrong, go and check.
Worth sorting the nine before any of the numbers, because it retires a lot of later hedging. Five are judgeable from the text alone — Forthrightness, Civility, Rigor, Specificity, Charisma. Three are relations between what someone said and an external record the pipeline doesn’t yet consult — Veracity, Strength, Authenticity — so they are the model’s estimate of plausibility, not a check. And Divination isn’t grounded at all: it scores whether a prediction is checkable and plausible when made, not whether it came true. Read the first group as measurement and the last four as an opinion with a quote attached.
Top and bottom. Left: the highest-ranked card with all nine attributes scored — Adrian Rurawhe, who was Speaker of the House until the end of 2023. Right: the lowest of the 130 ranked cards. The nine icons are the ones listed above, reading left-to-right and top-to-bottom in the same order; card rank is the geometric mean, so one bad attribute drags the whole card down. Which is also a flaw: cards are ranked on the geometric mean of whatever subset was scored, so an MP with a missing attribute — often their weakest — gets a free pass. Cards with eight attributes have a median rank of 45th; cards with all nine, 71st. That’s why I’ve shown Rurawhe here rather than the actual number-one card, and it needs fixing. Portraits from Wikimedia Commons.
The data behind it: 135 MPs who have sat in this Parliament, 45,477 statements scored 94,429 times — most statements bear on more than one attribute — spanning October 2023 – July 2026. About 69% comes from Hansard, the verbatim parliamentary record (190 sitting days of the current Parliament, ~7.5 million words), plus 2,976 party press releases and 65 post-Cabinet press conferences. Hansard matters most because it’s the only source that is adversarial, on the record, and covers every MP equally. Press releases tell you what a party wants you to think; Hansard tells you how they behave when someone is arguing back.
Nobody passes
Against a pass mark of 100, here is the headline.
Of the 1,150 MP-attribute scores in the dataset, not one reaches 80. The highest score any MP holds on anything is Adrian Rurawhe’s civility, at 79. The best whole card in Parliament is 66. The worst is 44.5.
Two caveats before you read anything into that, because they’re load-bearing and they belong here rather than at the bottom of the page. The scoring is not yet validated — the test set I evaluated against turns out to contain no New Zealand speech at all, which I explain below. And the displayed numbers are deliberately compressed: thin evidence is shrunk toward the population average, which is the right call for ranking but means the 79 ceiling is partly my estimator and not just Parliament. On raw unadjusted means, thirty-four MP-attribute scores clear 80, five of them on ten statements or more — Paulo Garcia and Greg Fleming both average 86 raw civility over thirteen statements. Individual statements do hit 100. The claim isn’t that no MP is ever exemplary; it’s that none of them sustains it, and that even ungenerously to my own headline, nobody is close to the standard.
Rigor is the bleakest of the nine. Of 14,503 statements scored for whether they argue from evidence rather than fallacy, three clear 90. None clear 95.
Civility averages 54, and about one in ten scored statements in the House lands at 20 or below — which is to say, is essentially just an insult. That’s roughly 1,115 of them across 190 sitting days, or six a day. I can’t tell you that’s the base rate in the chamber, because the model chooses which statements to surface and I haven’t measured whether it samples evenly — that’s the next thing I’d validate. What I can say is that the insults aren’t rare enough to need hunting for.
The league table on the site is therefore a bit misleading, and I want to say so plainly: it isn’t a ranking of good MPs and bad MPs. Everybody fails. The ranking is just the order they fail in. A civility average of 54 isn’t “middling” — it’s the class average on a test where the pass mark is 100.
Though I should own where that framing comes from, because it’s mine and not the instrument’s. The rubrics weren’t built as 0-fails/100-passes. Civility’s midpoint of 0.5 is anchored at “harsh but legitimate criticism of policies, ideas, or actions” — which is acceptable conduct, not failure — and the same is true of rigor and specificity. Its 0.9 exemplar is near-ideal parliamentary speech, so the rubric never really expected 1.0 either. If you want to argue that a 54 means “mostly legitimate criticism with a bad day in it”, the rubric is on your side and I’m not going to pretend otherwise. What survives the argument is the tail: one in ten statements at 20 or below is a personal attack by anyone’s rubric, and nobody’s average clearing 80 is a fact about the distribution rather than about where I put the pass mark.
Seven ways to fail
One example per attribute, all verbatim, all linked. The number is the score that statement received. These are the clearest illustrations I could find rather than a balanced sample — the site has 130 cards on the front page (every MP with at least six of the nine attributes scored) if you want to check your own team.
Forthrightness — 5/100. The Mitchell tautology from the top of this post. Asked to explain a term, he defines it with itself.
Specificity — 5/100. A whole sentence, committing to nothing:
“We are confident we will have effective legislation that will deliver good outcomes.” — Casey Costello, NZ First (Hansard, 27 Feb 2024)
Veracity — 5/100. The lowest veracity score in the dataset, and not a slip of the tongue — he builds two sentences on it. The Māui dolphin is a real, critically endangered subspecies of Hector’s dolphin:
“I may make a few remarks about a dolphin that doesn’t exist, otherwise known as the Māui dolphin. It was a contrivance, a fiction, put together by some underemployed academic down in the South Island.” — Shane Jones, NZ First (Hansard, 11 Feb 2025)
One caveat that belongs right here rather than at the bottom of the page: veracity is the attribute the model is least equipped to judge. Its own prompt tells it to score how well-evidenced a claim looks, not whether it is true — there is no fact-checking step wired in. I picked this example because the underlying fact needs no fact-checking step, but most veracity scores are plausibility judgements and should be read as such.
Rigor — 5/100. An ad hominem resting on a premise invented in the same breath (“they openly admit it”). Not in the House, this one — a party press release, where nobody is arguing back:
“The only thing most of them seem to be experienced at is being a bunch of Communists – they openly admit it.” — Winston Peters, NZ First (party press release, 28 Mar 2026)
Civility — 5/100. A jibe at another member’s mental capacity, mid-answer:
“And for the member who has just sprung into life—although perhaps not mentally—” — David Seymour, ACT (Hansard, 20 Aug 2025)
Divination — 10/100. A cheerful forecast that was implausible when he made it — the coalition’s position on the bill was already known, and National, ACT and NZ First duly voted it down 68–54 that afternoon. Worth being clear about what the score means: divination judges whether a prediction is checkable and plausible at the time of speaking, not whether it came true. The pipeline has no outcome-checking step. I confirmed the vote separately, from the division record.
“It’s a good bill, and I’m sure everyone in the House will support it—I certainly do.” — Duncan Webb, Labour (Hansard, 30 Jul 2025)
Authenticity — 30/100. My favourite of the lot, because both halves are from the same speech on the same afternoon. First:
“It’s the Greens—it’s the sanctimonious ‘tofu wouldn’t melt in your mouth’ Greens—who are the ones who are driving the division and acrimony within our body politic, aided and abetted by Labour.”
and then, minutes later:
“civility in politics and the ability for New Zealanders to talk to each other and to push back against that global direction of division and polarisation is an important part not only of our democracy but of how we succeed as a country.” — Paul Goldsmith, National (Hansard, 8 Oct 2025)
Two attributes are missing from that list, for two different reasons.
Charisma is missing because it is very nearly a duplicate of Civility. Its prompt asks whether a statement persuades across the aisle and says outright that crowd-pleasing attacks on opponents score low — so a charisma failure is usually the same insult, counted again. At the statement level the two correlate at 0.96. There was no charisma example to show you that wasn’t already up there under Civility, which is a real problem with the instrument and one I come back to below.
Strength is missing because of a bug I found while pulling these quotes. Strength and Authenticity are usually discussed rather than demonstrated — an MP stands up and describes an opponent’s broken promise — and my pipeline files that low score under the speaker. So MPs currently lose points for pointing out someone else’s failure. It affects about 37% of the low Authenticity examples and 7% of the low Strength ones, and it is directional: opposition MPs are hit 4.1 times as often as government ones on Authenticity (9.8% vs 2.4%) and 7.8 times as often on Strength (4.7% vs 0.6%), since scrutinising the government is their job. Authenticity stays on the list only because Goldsmith’s example is both halves about himself; Strength is out entirely.
It isn’t that they can’t
All of which would be unfair if the standard were unreachable. It isn’t.
Scroll the top of the civility scale and most of what you find is ceremony — tributes, condolences, farewells to retiring members, thanks to submitters. Warm, genuine, and carrying no adversarial stake at all. The chamber is entirely capable of civility in the part of the day where nothing is being contested.
The interesting cases are the ones where something is contested and it happens anyway. Chris Bishop, mid-argument, at 95: “the member does make a good point around funding pressure on the National Land Transport Fund. That is true.” (Hansard, 8 Apr 2025) Vanushi Walters — the top-ranked card — at 90, telling the other side “I do acknowledge to the members opposite that there are times when you do need to limit the right to freedom of expression.” (Hansard, 21 Aug 2025) Ginny Andersen, at 90 on the same day, granting that “demonstrations outside people’s homes can cross the line, and there are times when they can be deeply intrusive to families” (Hansard, 21 Aug 2025). None of these cost anything except the momentary pleasure of not conceding.
And then the single clearest illustration in the dataset, which is one MP, in one debate, on one afternoon. Simon Court, on the Resource Management Act, 31 March 2026 — scored 85:
“So the member Tamatha Paul makes a good point: the Resource Management Act is broken.”
and, in the same debate, to the same member — scored 20:
“what you Sim City obsessed planning nerds believe should be the right number of houses, Tamatha Paul” — Simon Court, ACT (Hansard, 31 Mar 2026)
He is plainly capable of the first. He was capable of it on the same afternoon, in the same debate, to the same person — and did the second anyway.
Two more for the record. One of the three statements that clear 90 for rigor, from a minister declining to spin a statistic running in his own favour:
“But we need to be careful not to draw simple conclusions from complex data. No single factor explains year-to-year changes in deaths and serious injuries, and it’s still too early to say whether this represents a long-term downward trend.” — Chris Bishop, National (speech to the AA annual conference, 27 Mar 2026)
And from the other side of the House, correcting a misused number with the source attached:
“The Prime Minister’s statements where he talks about ‘people on the youth benefit are now languishing on welfare for an average of 24 years’ — that’s also not backed up by the Taylor Fry report; that’s a forecast.” — Ricardo Menéndez March, Green (Hansard, 1 May 2024)
That’s the standard. MPs from every party reach it. It just isn’t where the mass of the distribution sits.
Where the failures cluster
Two patterns in the rankings that I didn’t expect, and can describe better than I can explain.
One caution before either of them. The nine attributes are not nine independent things. At the statement level Civility and Charisma correlate at 0.96, and Rigor with each of them at about 0.85 — and you can see why in the prompts, because all three penalise the same move, the personal attack. Civility marks down strawmanning; Rigor lists strawman and ad hominem first; Charisma says outright that crowd-pleasing attacks on opponents score low. So one jibe costs an MP a third of their card, and the geometric mean compounds it. Whatever the bottom of the table is measuring, it is measuring it about three times.
The party you’d blame isn’t the variable. I assumed the data would sort the parties. It doesn’t really — the party averages sit within a few points of each other, and MPs vary far more within each party than the parties do between themselves. Whichever team you came here to have confirmed, it won’t be: the top five cards are three Labour MPs, a National MP and an NZ First MP.
That’s true of the means, and it is the honest headline. It’s less true of the tails, and since the tails are what I’ve been naming people over, I should show them. Of the top twenty cards, eleven are Labour and five National; of the bottom twenty, five are National, four are Te Pāti Māori, three each Green, ACT and NZ First, two Labour. Te Pāti Māori have four of their seven MPs in the bottom twenty and none in the top. I don’t read that as a finding about honesty. I read it as the point above: an instrument that triple-counts adversarial register will rank the most rhetorically confrontational parties lowest, and a party of seven MPs whose entire role is confrontation with the Crown is going to sit badly with a rubric that rewards the technocratic voice. That’s a limit of what I’ve built, not a verdict on them.
The bottom of the table is the leadership. Three cards sit clearly below the rest: Winston Peters (44.5), Shane Jones (44.6) and Rawiri Waititi (48.1). After that the table compresses so tightly that individual ranks stop meaning much — twenty-five places separated by three points, against per-attribute confidence intervals of ±4 — so read the band, not the number. What the band says is that the last ten places hold the Prime Minister, a Cabinet minister, and five party leaders and co-leaders; Chris Hipkins (82nd) and Chlöe Swarbrick (89th) are in the bottom half. The top of the table is backbenchers, a former Speaker, and select committee workhorses — mostly people you’d struggle to name.
It’s general rather than anecdotal, with one caveat: score correlates negatively with how much an MP talks (r = −0.25), though some of that is my own estimator, which pulls thinly-covered MPs toward the average. Among the twenty-one most-covered MPs the correlation weakens to about −0.10. The blunter fact survives — the twenty most-quoted MPs average 87th of 130.
I don’t know why, and I’d rather flag the guesses as guesses. Part of it is certainly selection — ministers front oral questions, which is the most adversarial hour in the building, so they’re sampled at their worst. Part of it might be that conduct like this is rewarded: a good select committee question travels nowhere, and Seymour’s funeral-parlour line travels a long way. Every one of these nine attributes is a virtue when you’re persuading the person in front of you and a liability when you’re playing to an audience. Whether MPs are actually responding to that, or whether the people who rise to lead parties were always like this, I can’t tell from this data. (Checking how often MPs cross party lines — voted.nz collects the divisions — would be one way in.)
What I can say is the plain version: the people whose words carry furthest, whom we’ve promoted highest, are the ones scoring worst. The visible part of our politics is the worst part of it.
The comparison isn’t fair yet
I’ve been ranking MPs against each other for a thousand words now, and I owe you the reasons that ranking isn’t yet a fair fight. There are three, and they all point the same way.
Where an MP’s words come from changes their score more than who they are. Scores differ enormously by source. A statement in a party press release scores 17.6 points higher on Forthrightness than one in Hansard, and 9.3 lower on Specificity, 7.2 lower on Civility, 6.1 lower on Rigor. The whole league table spans 21.5 points, from 44.5 to 66 — so a single source gap is most of the range I’ve been ranking people over. And the source mix differs enormously by party: 45% of ACT’s statements come from party releases, against 17% of NZ First’s and 18% of National’s. Post-Cabinet press conferences are government-only by construction — 5,269 statements from National MPs, 126 from NZ First and 24 from ACT, and none at all from Labour, the Greens or Te Pāti Māori. I used to be able to say this confound didn’t apply, because an earlier version of the dataset was Hansard-only and every MP was measured in the same adversarial chamber. That stopped being true when I added 23,935 press-release statements, and nothing has replaced the mitigation.
Forthrightness is close to a minister-only measure. The prompt scores instances where a politician is asked a question and either answers it or doesn’t — which in practice means oral questions, which in practice means ministers. Only 104 of the 135 MPs have a Forthrightness score at all, and it has by far the largest source gap of the nine. Comparing a minister’s Forthrightness to a backbencher’s is comparing two different tests.
Three of the nine structurally favour the executive. Specificity rewards “figures, mechanisms, timeframes, or named policies” — which is what a minister announcing a programme produces and what an opposition MP asking why it hasn’t happened does not. Strength measures delivery, and only ministers and members in charge of a bill can deliver anything; my own notes on the vote data say this plainly, that scoring an opposition backbencher low for not delivering repeats the attribution bug in a new form. Divination, as above, is measuring almost nothing.
None of that makes the distribution finding go away — nobody clearing 80 on anything is a fact about the whole chamber, whatever mix of sources it came from. But it does mean that the difference between 110th and 120th on the league table is not yet a fact about those two MPs, and I’d rather say so than let the ranking imply a precision it hasn’t earned.
What it isn’t
A prototype, and I’d rather undersell it.
The scores are automated estimates, not verdicts — and I want to be precise about how weak the validation currently is, because I got this wrong when I first published.
The evaluation I built runs against fourteen hand-labelled statements per attribute, and those statements are invented: fictional American council-meeting speech, “Senator Armstrong” and “Governor Miller”, not a line of Hansard in them. On that set civility scored r = 0.88 and veracity r = 0.73, which sounds reassuring and isn’t — it measures the model on a task that is not this one. NZ parliamentary speech is a different register and far more adversarial. A replacement set sampled from the real corpus exists but is not yet labelled. Until it is, treat every accuracy figure here — including the two I just quoted — as unvalidated. There is also only one labeller, so I have no inter-rater agreement to report, and some of the prompts carry few-shot examples I haven’t yet verified are disjoint from the test sets.
Accuracy still varies by attribute, predictably, and the prompts say so themselves. Civility and specificity are judgeable from the text alone. Veracity is not: deciding whether a claim about someone’s voting record is true means going and checking the voting record, which the model can’t do — those scores reflect plausibility, and should be read that way. Divination is worse: the prompt scores whether a prediction is falsifiable and plausible at the time of speaking, not whether it came true, and it shows. Across 126 cards every MP lands between 43 and 54 — an eleven-point spread for the whole Parliament, when a typical single card’s own confidence interval on divination is ±4 — which is about what you’d expect from a model asked to rate plausibility with no outcomes to check against. It shouldn’t be one of the nine terms in the geometric mean until it’s grounded, and it currently is. Strength and Authenticity carry the attribution bug described above and should be discounted heavily until I fix it.
There’s a subtler problem with reading verbatim speech literally. Hansard is lightly-edited spoken language, full of inversions, self-corrections and half-finished sentences, and a scorer that sees one statement with no surrounding discourse takes all of them at face value. A word-order slip reads as a false claim. I’ve seen enough of these in the low-veracity examples to say plainly: a single low score on a single sentence is not evidence of anything, and the reason every number links to its quote is so you can see which kind you’re looking at.
And coverage is uneven. A minister who fronts oral questions every week generates far more text than a first-term backbencher, so confidence differs enormously between cards; the site shows a confidence level per attribute for that reason.
Nothing here is an accusation of wrongdoing. It’s a starting point for looking, and the evidence is one click away.
Why bother
Measurement won’t fix politics. But notice how much we already measure about our politicians — polling, preferred PM, seat projections — and how little of it is about whether they’re doing the job well. We measure popularity with great precision and conduct not at all.
Which is a shame, because conduct is the one thing we could actually agree on. You and I can disagree completely about tax and still both want the person arguing our side to answer the question, get their facts right, and skip the jokes about hair. That’s a standard with no political side to it — and right now, on the public record, in our name, not one of the 135 people who have sat in this Parliament is meeting it.
There’s an election in November. It would be nice to be able to ask.
Have a look: nz-politician-scorecards.onrender.com — cards on the front page, the dataset for full stats and downloads, about for method and caveats. It’s on free hosting, so the first visit after a quiet period takes a few seconds to wake up.
Corrections, 7 August 2026. This post was revised after a review checked it against the project repo. The substantive changes:
- The original said the scores were validated against “small human-labelled test sets” at r = 0.88 and r = 0.73, and hedged only on sample size. Those test sets turn out to be invented American statements with no New Zealand speech in them, so the figures don’t measure this task at all. Rewritten in What it isn’t.
- The original described the card number as “just the mean of the evidence”. It’s a shrunk mean — thin evidence is pulled toward the population average — which is a deliberate design choice but partly explains the 79 ceiling. Noted where the ceiling is claimed.
- The original said Divination measures whether predictions come true. It doesn’t; it scores whether a prediction is checkable and plausible when made. The Duncan Webb example has been reframed.
- The original Veracity example was Willie Jackson, scored 10/100 on the sentence “fifty percent of Māori are in jails at the moment”. In context that is a word-order slip in live speech for “fifty percent of people in jails are Māori” — the sentences either side of it are about Māori over-representation in Corrections. Publishing it under a “Veracity — 10/100” heading, scored by an ungrounded attribute, was unfair to him and I’ve removed it. My apologies.
- Counts corrected: 45,477 distinct statements (94,429 is the number of statement-attribute scores, not statements); 190 sitting days of this Parliament, not 204; 130 cards on the site’s front page, not 135; “135 people we elected” → 135 who have sat in this Parliament, since New Zealand elected 123.
- The Simon Court pair is stamped to the same date but not the same hour; “an hour before” removed.
- Charisma, not Authenticity, was the second attribute missing from “Seven ways to fail”. Both attribution-bug multipliers now given separately.
- New sections added on attribute correlation, rubric framing, and why the comparison isn’t fair yet. Ranks in the bottom of the table are now given as a band rather than exact places.
The site’s about page carried the same stale validation claim and has been corrected too.