Every prediction is logged the moment it's made and graded when the game ends, wins and losses alike. Regular season and playoffs are reported separately; they are different games and we never mix them.
Grading starts with the new season.
The season is under way; the first graded calls land as games settle.
College coverage here is scores and schedules.
Each prediction is saved the moment the model makes it, before the game starts, and can't be changed afterward. When the game ends we grade it against the final score. Nothing gets quietly dropped: the record you see includes every settled call.
The probability table is the harder test. Being right 70% of the time is good; saying 70% and being right 70% of the time is what makes a probability worth trusting. That's what the "we said / it happened" columns compare.