On 9 August 2026 we deleted a prediction from CompactStats that was winning three times in four. It filled 44% of every fixture list and graded 102 from 136. Removing it made our published record worse. It was still the right call, and this is the arithmetic behind that.
Over 3.5 goals — four or more in a match — happens in roughly 30% of football. So under 3.5 happens in roughly 70%. Any model that says "under" on a game is starting from a 70% base rate it did nothing to earn: the coin is weighted before you speak. Our under-calls scored 75%. Against that baseline, the entire contribution of the model was about five percentage points, and at n=136 even that is inside the range chance produces on its own.
Contrast the other direction. Our Over 3.5 calls fire only when the rating clears 45% — 6% of fixtures — and they have landed 11 of 19, about 58%, against a 30% base rate. That is nearly double the baseline. It is a smaller, rarer, less flattering-looking record, and it is the one that means something.
A visitor reading "Under 3.5 — 102/136" reasonably concludes we are good at predicting low-scoring games. What they actually saw was football's default outcome, restated. The number was true and the impression it created was false, which is the definition of a misleading statistic.
There is a broader test buried here that we now apply to every call the site makes: compare to the naive baseline, not to zero. A 75% strike rate is excellent for a coin-flip market and mediocre for a 70% one. Betting sites quote raw strike rates precisely because the baseline is invisible to most readers. We would rather quote the lift.
Every Over 3.5 rating below 45% now shows the plain probability of four or more goals — "O 34%" — with no pick and no grading. The column is quieter and considerably less impressive-looking. It is also the honest description of what a model can tell you about a 30% event.
We applied the deletion to our frozen history too, so the winning calls disappear from the past record as well as the future one. If you find that odd, consider the alternative: keeping a 75% badge on the site while knowing what it was worth.
Method: walk-forward testing — every prediction is made using only information available before kick-off, then graded against what happened. Studies run June-July 2026 on the site's own dataset. Questions: contact us.