Findings
What we measured in this project, and how the measurements changed our product decisions.
1. The model does not beat the market
Walk-forward evaluation over 8 leagues and roughly 19,000 matches. In none of the 8,434 predictions was the model more accurate than the bookmakers' closing line. Per-league table.
This result set the product's positioning: we cannot claim "winning tips", because the measurement does not allow it.
2. Where we lose: not calibration, but resolution
A Murphy decomposition shows most of the loss comes from resolution. The percentages we state are honest; we simply cannot discriminate as sharply as the market.
3. Post-hoc calibration does not help — and that is a trap
It appears to improve things in-sample and loses out-of-sample. That is why it stays off without holdout evidence.
4. A prior for newly promoted teams
A promoted side has no top-flight history; treating it as league-average is too optimistic. A measured prior improved all eight leagues.
Data and coverage
43,937 fixtures, 8,434 walk-forward predictions, 8 leagues. Source: football-data.co.uk. Raw data is not redistributed; only derived output is published here.