Every tipster publishes a win rate. A win rate cannot tell a good model from a lucky one, and it cannot tell a model that says 60% and is right 60% of the time from one that says 90% and is right 60% of the time. These are the numbers that can.
Scored as a forecast
ScooreCast model — Brier 0.6193, log loss 1.0315, top pick right 48.8%
Always pick home — Brier 1.1382, log loss 19.6564, top pick right 43.1%
No information (a flat 33/33/33) — Brier 0.6667, log loss 1.0986, top pick right 43.1%
The model's probabilities carry information: log loss 1.0315 against 1.0986 for no information at all.
Reliability — when the model says 45%, does it happen 45% of the time?
Model said 0–20%: 455 outcomes, mean stated 14.9%, happened 18.5%
Model said 20–40%: 2345 outcomes, mean stated 28.5%, happened 28.0%
Model said 40–60%: 894 outcomes, mean stated 47.9%, happened 47.8%
Model said 60–80%: 170 outcomes, mean stated 66.8%, happened 65.9%
Model said 80–100%: 21 outcomes, mean stated 85.4%, happened 71.4%
A well calibrated model has those two figures close together in every row. Rows counting only a handful of outcomes say nothing at all.
What the confidence stars were worth
3 stars: 63 of 134 came off, 47.0% — the tips claimed 49.7% on average.
4 stars: 178 of 323 came off, 55.1% — the tips claimed 56.7% on average.
5 stars: 83 of 119 came off, 69.7% — the tips claimed 73.0% on average.
The rating decides whether a tip is published at all, so what it delivered belongs on this page. On the present sample the bands cannot be separated: three and four stars differ by under a point and each rate carries a margin of error more than twenty points wide. That is not evidence the rating is worthless, only that there is not yet enough settled football to judge it.
What is in the sample
719 of 1295 were void (below the publication threshold). They are scored here but excluded from the public record, by design — they are the close fixtures, and dropping them would bias calibration toward the confident end.
How to read this honestly
Fixtures held back as too close to call are included here and excluded from the public win rate. Dropping them would bias calibration toward the confident end and flatter the model.
Every tip is scored with the probabilities it carried when it was published. Tips freeze at kickoff and are never edited.
The model leans heavily on market odds today. Beating a de-vigged closing line is a far harder bar than beating no information, and this page will report that comparison once the sample can support it.
Club and competition names and crests belong to their respective owners and are used only to identify them. ScooreCast is not affiliated with or endorsed by any club, league, federation or ESPN.