What we learned auditing 104 leagues

A reasonable worry about a model covering 150 competitions is that it works in the Premier League and flounders in the Icelandic second tier. We ran the test: every covered league, walk-forward, all four markets, asking one question — does the model beat simply quoting that league's average?

The result nobody wants to publish

Across 104 leagues with enough history to judge, 24 came back with the model performing slightly worse than the league average. That sounds alarming until you look at the sizes: those leagues hold 200–700 graded matches each, and the gaps are two or three thousandths of a Brier score. With 104 leagues tested, a couple of dozen landing slightly negative is precisely what chance produces around a positive mean. None came close to the standard set by our qualifier audit, where 1,150 games gave a flat coin-flip verdict and we stopped making picks.

The one real pattern

Of eleven leagues where the result model failed to beat "always pick the home side", seven were third tier or lower — Scottish League Two, Brazil Serie C, Portugal's Liga 2, Argentina's second tier, J3, USL, Mexico's Expansión. Individually each is within noise. As a cluster it is suggestive: in chaotic divisions, team-strength estimates are mostly noise, and the model's away picks burn value that a naive home bias would keep.

What we did about it: nothing, yet

Every restriction this site ships was built on evidence at the thousand-game scale. Gating an entire division on 300 games of ambiguity would be the opposite of that discipline — and a model tuned to whichever leagues looked bad last month is a model fitted to noise. Instead, the frozen pre-match ledger we started keeping on 31 July now records every published pick, per league, immune to hindsight. In September we will re-run this audit on real shipped picks rather than reconstructions, and if the lower-division weakness survives contact with clean data, the restriction writes itself.

The uncomfortable version of scientific honesty is not "we found a problem and fixed it". It is "we found something that might be a problem, and the responsible response is to wait for better evidence".

Method: walk-forward testing — every prediction is made using only information available before kick-off, then graded against what happened. Studies run June-July 2026 on the site's own dataset. Questions: contact us.