Walk-Forward Evaluation¶
elote.walk_forward() predicts each period from everything before it, then learns that
period. The returned elote.WalkForwardReport carries aggregate accuracy, log loss and
Brier score, and a reliability table showing which probability ranges are overconfident.
Reliability tables¶
report.reliability is a tuple of elote.ReliabilityBin records, one per
equal-width bin on [0, 1]. Choose the number of bins with calibration_bins (default 10; a
positive integer, booleans and non-integers raise ValueError).
Bins are left-inclusive and right-exclusive, except the final bin, which also contains 1.0.
Each prediction is binned before the log-loss clamp is applied.
The population is exactly the
report.predictionspopulation: decisive bouts between competitors already seen, after thewarmupperiods. Draws, missing outcomes, unseen competitors and warmup periods are excluded, so bin counts sum toreport.predictions.mean_predictedis the mean probability that the first side wins;observed_rateis the fraction of those bouts the first side won. Empty bins havecount == 0andNonefor both.Only bin aggregates are stored. This is an empirical expected-score calibration of decisive bouts, not a three-outcome draw-probability model.
from elote import EloCompetitor, group_by_period, walk_forward
periods = group_by_period(rows)
report = walk_forward(EloCompetitor, periods, warmup=2, calibration_bins=5)
for b in report.reliability:
if b.count:
print(f"[{b.lower:.1f}, {b.upper:.1f}) n={b.count} "
f"predicted={b.mean_predicted:.3f} observed={b.observed_rate:.3f}")
A bin whose mean_predicted is well above its observed_rate is overconfident.