Accuracy

Every prediction is logged the moment it's made and graded when the game ends, wins and losses alike. Regular season and playoffs are reported separately; they are different games and we never mix them.

NBA

Every game the model called, all season
75.0% correct · 252 settled calls · regular season
Playoffs, kept separate: 44.0% on 25 calls
Spread 46.0% 163 calls
Totals 53.4% 146 calls
Season accuracy as it settled, day by day
coin flip · 50%
Do our probabilities mean it? Group the calls by how confident we were, then check each group against what happened
We saidGamesWe said (avg)It happened
Under 45% 89 29% 29%
45 to 55% 22 51% 59%
55 to 65% 34 60% 77%
65 to 75% 31 71% 74%
75% and up 76 83% 84%

MLB

Every game the model called this season, graded daily
54.6% correct · 1501 settled calls · regular season
Spread 50.9% 641 calls
Totals 50.5% 658 calls
Season accuracy as it settled, day by day
coin flip · 50%
Do our probabilities mean it? Group the calls by how confident we were, then check each group against what happened
We saidGamesWe said (avg)It happened
Under 55% 742 52% 52%
55 to 65% 725 59% 57%
65 to 75% 34 67% 68%

NHL

Grading starts with the new season.

Soccer

Every league match the model called; a 3-way game, so 33% is chance
50.5% correct · 436 settled calls · regular season
Season accuracy as it settled, day by day
chance · 33%
Do our probabilities mean it? Group the calls by how confident we were, then check each group against what happened
We saidGamesWe said (avg)It happened
Under 55% 341 44% 45%
55 to 65% 64 59% 72%
65 to 75% 25 69% 64%
75 to 85% 6 78% 83%

NFL

The season is under way; the first graded calls land as games settle.

NCAA

College coverage here is scores and schedules.

How the grading works

Each prediction is saved the moment the model makes it, before the game starts, and can't be changed afterward. When the game ends we grade it against the final score. Nothing gets quietly dropped: the record you see includes every settled call.

The probability table is the harder test. Being right 70% of the time is good; saying 70% and being right 70% of the time is what makes a probability worth trusting. That's what the "we said / it happened" columns compare.