An account of a mistake in our own method, and the checks that now catch it. Brand names, categories, competitors and markets have been removed.
The finding that looked decisive
A read came back with a clean story. On questions phrased as superlatives, the best, the biggest, the most, the brand was named in 83% and 85% of answers. On questions phrased around fit, which one suits someone like me, the same brand managed 29% and 38%.
Better still, the brand's score swung up by 54 points between the two phrasings while every competitor's swung down. That reads like a real asset: this is a brand people reach for when they want the top of the category.
It was not.
What we found when we checked the questions
Someone asked which superlatives had actually been put to the engines. Six had. Five of them were records the brand holds by definition, facts about the brand that no competitor could have been named for, because they are not true of anyone else.
The sixth was a superlative the brand does not hold. On that one it was named 3 times out of 30, against a competitor's 27.
The finding was not a strength. It was the question set describing itself back to us.
The control costs one question
The fix is small, and it has to happen before the read runs. When a question set is written, mark every question that only one brand could win, and make sure the set contains superlatives the brand does not hold alongside the ones it does.
Afterwards is too late. A real strength and a definitional one produce exactly the same shape in the results, so they cannot be separated by looking at the output. The information needed to tell them apart is in how the question was written, and if it was not recorded then, it is gone.
A related trap runs the other way. A question that names only your competitors will score you a zero every time. Leave a few of those in a set and they will drag a score down for a reason nobody can see later.
Panels catch arguments. Only a machine catches numbers.
We put findings in front of a review panel before they ship, and panels are good at what they are good at. They kill weak reasoning, they ask who the reader is, and they spot an argument that flatters us.
They do not catch arithmetic. In one build, two figures that had been keyed in by hand went past a five-person panel without a flicker, then failed on the first automated run against the underlying data:
- A paired count printed as 0 of 30. The data said 5 of 30.
- A claim that the brand came first in all four groups. It was second in one of them, 63 against 69.
Both sounded right, both were wrong, and both were on their way to a client. Every number in a report is now generated from the frozen data rather than typed, and the check that enforces it runs on every build.
A ranked list is not a rank order
One more, because it is the easiest mistake to make and the hardest to see. When ten options are scored and the top five sit within 1.0 and 0.6 points of each other, that is a band, not an order. Reporting it as first, second and third invents a precision the sample cannot carry, particularly when each option rests on four to six questions.
A dead heat gets reported as a dead heat, with a plain statement of which options are genuinely separated and which are not.
Why publish this
Because the failure is not unusual, and every one of these mistakes flatters the person selling the report. A method is only worth trusting if it is built to catch its own good news.