Chronological tuning and final evaluation¶
Reserve later periods before searching for parameters. Here we use six local, deterministic weekly periods with point scores, comparing small Elo and Pythagorean grids. The first four periods are development data; the last two remain untouched by selection. This illustrative schedule establishes no universally best rating system.
Run the example from a checkout with the base package installed:
uv run --no-sync python examples/chronological_selection.py
No dataset downloads, plotting packages or optional extras are needed.
The executable source¶
"""Select on development periods, then evaluate once on later periods, offline."""
from datetime import datetime, timedelta, timezone
from elote import EloCompetitor, PythagoreanCompetitor, group_by_period, tune, walk_forward
INITIAL_WARMUP = 1
DEVELOPMENT_PERIODS = 4
SCORE_KEYS = ("home_points", "away_points")
SYSTEMS = {
"Elo": (EloCompetitor, {"k_factor": [16, 32]}),
"Pythagorean": (PythagoreanCompetitor, {"exponent": [1.5, 2.37]}),
}
def scored_schedule():
"""Six weekly periods; the last two are reserved before any search."""
games = [
[("A", "B", 24, 14), ("C", "D", 21, 17)],
[("A", "C", 28, 20), ("B", "D", 17, 10)],
[("A", "D", 21, 14), ("B", "C", 14, 14)],
[("A", "B", 20, 17), ("C", "D", 24, 10)],
[("A", "C", 17, 24), ("B", "D", 21, 21), ("E", "A", 10, 20)],
[("A", "D", 28, 14), ("B", "C", 17, 20), ("E", "B", 14, 21)],
]
start = datetime(2024, 1, 1, tzinfo=timezone.utc)
return [
(
a,
b,
1.0 if home > away else 0.0 if home < away else 0.5,
start + timedelta(weeks=week),
{SCORE_KEYS[0]: home, SCORE_KEYS[1]: away},
)
for week, period in enumerate(games)
for a, b, home, away in period
]
def select_and_evaluate(rows):
"""Return development candidates, frozen choice, and a separate final report."""
periods = group_by_period(rows)
development = periods[:DEVELOPMENT_PERIODS]
candidates = []
for name, (system, grid) in SYSTEMS.items():
results = tune(
system,
grid,
development,
metric="log_loss",
warmup=INITIAL_WARMUP,
score_keys=SCORE_KEYS,
)
candidates.extend((name, result) for result in results)
candidates.sort(key=lambda candidate: candidate[1].report.log_loss)
name, selected = candidates[0]
system = SYSTEMS[name][0]
final_report = walk_forward(
system,
periods,
competitor_params=dict(selected.params),
warmup=len(development),
score_keys=SCORE_KEYS,
)
return candidates, name, selected, final_report
def print_report(label, report, *, boundary, warmup):
print(
f"{label}: protocol=predict-then-learn; {boundary}; warmup={warmup}; "
f"predictions={report.predictions}, skipped={report.skipped}, draws={report.draws}; "
f"log_loss={report.log_loss:.6f}, brier={report.brier:.6f}, accuracy={report.accuracy:.6f}"
)
def main():
rows = scored_schedule()
periods = group_by_period(rows)
for index, period in enumerate(periods):
print(f"Period {index}: {period[0][3].date()} through {period[-1][3].date()}, rows={len(period)}")
candidates, name, selected, final_report = select_and_evaluate(rows)
for candidate_name, candidate in candidates:
print_report(
f"Development selection {candidate_name} {candidate.params}",
candidate.report,
boundary="fit period 0; score periods 1-3",
warmup=INITIAL_WARMUP,
)
print(f"Frozen choice: {name} {selected.params} (minimum development log loss)")
print_report(
"Final evaluation",
final_report,
boundary="replay periods 0-3; score untouched periods 4-5",
warmup=DEVELOPMENT_PERIODS,
)
print("Parameters stay fixed; ratings learn each holdout period after all its predictions.")
if __name__ == "__main__":
main()
Reading the protocol¶
group_by_period orders the timestamped rows into ISO weeks. Tuning uses only
periods 0-3 and an explicit one-period initial warmup. Each grid candidate scores
five decisive development bouts. tune ranks configurations by development
log loss; the example then chooses across both systems using that same metric.
These are selection scores, reused to make the choice, not independent final
evidence. Class parameters belong in the grids; constructor options, if needed,
would be supplied separately through base_competitor_kwargs.
Once selected, the system and parameters are fixed. The final walk_forward
starts from fresh ratings, replays the complete schedule, and uses all four
development periods as warmup. Only periods 4-5 contribute to the final metrics.
Using the initial tuning warmup here would incorrectly include development
predictions in the final report.
Every prediction within a period uses ratings from earlier periods: no result in that period informs another prediction in the same period. After scoring, the whole period is learned. This also happens on the final holdout: parameters stay fixed while ratings evolve, so this is adaptive predict-then-learn evaluation. It differs from predicting an entire holdout with frozen ratings. The holdout is untouched by search, not withheld forever from rating updates.
score_keys forwards the custom point-score attributes through tuning and
final replay. Pythagorean uses those points; the default Elo configuration here
learns win/loss/draw outcomes without margin scaling.
The final population is exactly four predictions, one skipped bout and one draw. Team E first appears in period 4, so its decisive bout is skipped because it has no prior rating; it is learned and can be scored in period 5. Draws are also learned but excluded from accuracy, log loss and Brier. Warmup bouts contribute to fitting, not to these reported counts. The output prints period dates, scoring boundaries, warmup, protocol and population beside every metric.
Keep the final report separate from selection scores, and do not revise the choice after inspecting it. If you change the search based on final results, reserve new later data for another independent evaluation.