Premier League β Validation
Every played match: our pre-match forecast vs. what actually happened. The receipts. π§Ύ
Wait β how can 100.0% accuracy still beat a coin-flip on RPS?
Because accuracy and RPS measure different things, on different scales β they're not in competition.
- Accuracy only asks: was our single top pick right? It's blind to the probabilities, and a model almost never makes draw its top pick β so a draw-heavy round drags it down even when the forecast was good.
- RPS grades the whole probability spread (e.g. 55 / 25 / 20) against a uniform βcoin-flipβ (33 / 33 / 33). The coin-flip never commits, so it's never punished hard β which makes it a surprisingly tough baseline, and means a few % of skill is a genuine edge.
Our committed forecast cuts both ways (lower RPS = better):
| Match | Our forecast | Result | Our RPS | Coin-flip RPS |
|---|---|---|---|---|
| Favourite delivers | 55 / 25 / 20 | home win | 0.12 β | 0.28 |
| We back home, it's a draw | 55 / 25 / 20 | draw | 0.17 | 0.11 β |
When a favourite wins we crush the baseline; on a draw, the hedging coin-flip actually beats us. Matchday 1 had 9 draws in 24 games, so the wins and losses partly cancel β the net edge is the βskillβ figure above. A few % over uniform is roughly what bookmakers manage, so it's a real, healthy edge.
Matchday 1
| Match | Forecast (1 / X / 2) | Our call | Predicted | Actual | |
|---|---|---|---|---|---|
| 77.3 / 14.8 / 7.9 | Arsenal | 2β0 | 3β0 | ✓ |
✓ = our most-likely result was correct Β· ⚛ = exact scoreline nailed Β· multiclass Brier score 0.08. Click any match for its full forecast breakdown.