Evaluation API Reference¶
Walk-forward evaluation¶
benchmark_competitors compares systems on one frozen train/test split. Walk-forward
evaluation predicts each period from everything before it and then learns that period, and it
reports log loss and Brier score alongside accuracy. compare_walk_forward runs that
protocol for several systems over the same periods and warmup, so the numbers are comparable.
from elote import (
EloCompetitor, GlickoBoostCompetitor, SyntheticDataset,
compare_walk_forward, group_by_period,
)
periods = group_by_period(SyntheticDataset(num_competitors=20, num_matchups=1000, seed=1).load())
result = compare_walk_forward(
{
"elo": EloCompetitor,
"elo_k32": (EloCompetitor, {"competitor_params": {"k_factor": 32}}),
"glicko_boost": GlickoBoostCompetitor,
},
periods,
warmup=4,
)
print(result)
result.ranking()[0]["system"] # best log loss
Each system is a class, or a (class, options) pair where options may hold
competitor_params and base_competitor_kwargs. Invalid systems raise before anything is
evaluated. Walk-forward numbers are not interchangeable with frozen-split numbers: they score
a different population of bouts, and warmup decides how much early history is excluded.
same_population is False when systems scored different bout counts.
Walk-forward evaluation and hyperparameter search for rating systems.
evaluate_competitor() trains on one split and then predicts a held-out
split with frozen ratings. That answers “how well do these ratings survive going stale”,
which is a real question but rarely the one being asked. The usual question is how a system
performs in the way it would actually be used: predict the next round of results from
everything that has happened so far, then fold those results in and step forward.
This module provides that protocol, the metrics that can see a system’s calibration as well as its picks, and a grid search over competitor parameters.
- Example:
>>> from elote import EloCompetitor, walk_forward, group_by_period >>> periods = group_by_period(rows) >>> report = walk_forward(EloCompetitor, periods) >>> report.accuracy, report.log_loss
- class elote.evaluation.ReliabilityBin(lower: float, upper: float, count: int, mean_predicted: float | None, observed_rate: float | None)[source]¶
Bases:
objectOne equal-width probability bin of a reliability table.
The bin covers
[lower, upper); the final bin of a table also includes 1.0.- Attributes:
lower: Inclusive lower bound. upper: Exclusive upper bound (inclusive for the final bin). count: Scored predictions whose probability fell in the bin. mean_predicted: Mean predicted probability that the first side wins, or
Noneif empty. observed_rate: Fraction of the bin’s bouts the first side actually won, orNoneif empty.
- count: int¶
- lower: float¶
- mean_predicted: float | None¶
- observed_rate: float | None¶
- upper: float¶
- class elote.evaluation.TuningResult(params: Dict[str, Any], report: WalkForwardReport)[source]¶
Bases:
objectOne point of a
tune()grid search.- params: Dict[str, Any]¶
- report: WalkForwardReport¶
- class elote.evaluation.WalkForwardComparison(reports: Dict[str, WalkForwardReport], warmup: int, periods: int, same_population: bool)[source]¶
Bases:
objectWalk-forward reports for several systems run over the same periods and warmup.
- Attributes:
reports: Label to
WalkForwardReport, in the order the systems were given. warmup: Leading periods used for fitting but not scored, shared by every system. periods: Number of periods every system was run over. same_population:Truewhen every system scored the same number of bouts andskipped and drew the same numbers. When
Falsethe figures describe different row populations and should not be compared as-is.
- periods: int¶
- ranking() List[Dict[str, Any]][source]¶
Rows sorted by log loss, best first; systems with no scored bouts (NaN) come last.
Every row carries the protocol (
warmup,periods) and the scored-row counts.
- reports: Dict[str, WalkForwardReport]¶
- same_population: bool¶
- warmup: int¶
- class elote.evaluation.WalkForwardReport(predictions: int, skipped: int, draws: int, accuracy: float, log_loss: float, brier: float, by_period: Tuple[Tuple[int, int, float], ...] = (), reliability: Tuple[ReliabilityBin, ...] = ())[source]¶
Bases:
objectMetrics from a walk-forward run.
- Attributes:
predictions: Bouts that were both scored and predictable. skipped: Bouts skipped because a competitor had not been seen yet. draws: Drawn bouts, excluded from every metric below. accuracy: Fraction of predictions on the correct side of 0.5. log_loss: Mean negative log likelihood. Sees calibration; accuracy does not. brier: Mean squared error of the predicted probability. by_period:
(period_index, predictions, accuracy)per scored period. reliability: Equal-widthReliabilityBinrecords over exactly thepredictionspopulation (decisive, predictable, post-warmup bouts; draws are excluded). Binned on the original prediction, before the log-loss clamp.
- accuracy: float¶
- brier: float¶
- by_period: Tuple[Tuple[int, int, float], ...] = ()¶
- draws: int¶
- log_loss: float¶
- predictions: int¶
- reliability: Tuple[ReliabilityBin, ...] = ()¶
- skipped: int¶
- elote.evaluation.compare_walk_forward(systems: Dict[str, Any], periods: Sequence[Sequence[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]], *, warmup: int = 0, score_keys: Tuple[str, str] | None = None) WalkForwardComparison[source]¶
Run
walk_forward()for each system on the same periods and warmup.- Args:
- systems: Label to a competitor class, or to a
(class, options)pair where optionsmay holdcompetitor_paramsandbase_competitor_kwargsexactly aswalk_forward()takes them.
periods: Ordered periods of dataset rows, as produced by
group_by_period(). warmup: Leading periods used for fitting but not scored, applied to every system. score_keys:(a_score_key, b_score_key)naming each row’s two point scores.- systems: Label to a competitor class, or to a
- Returns:
WalkForwardComparison: Each system’s report, a log-loss ranking and a printable table.
- Raises:
- InvalidParameterException: If
systemsis empty or any entry is malformed, is not a competitor class, or names a parameter its class does not have. Raised before any system is evaluated.
ValueError: If
warmupis negative or not smaller than the number of periods.- InvalidParameterException: If
- elote.evaluation.group_by_period(rows: Iterable[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]], key: Callable[[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]], Any] | None = None) List[List[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]][source]¶
Group dataset rows into chronologically ordered periods.
A period is the unit of “predict, then learn”: everything inside one is predicted before any of it is used for fitting, which is what stops a result informing a bet placed on the same afternoon.
- Args:
rows: Dataset rows, in any order. key: Maps a row to its period. Defaults to the ISO calendar week of the row’s
timestamp, which suits weekly league sports. Rows without a usable timestamp are collected into one leading period.
- Returns:
A list of periods, each a list of rows, ordered by period key.
- elote.evaluation.tune(competitor_class: Type[BaseCompetitor], param_grid: Dict[str, Sequence[Any]], periods: Sequence[Sequence[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]], *, metric: str = 'log_loss', **walk_forward_kwargs: Any) List[TuningResult][source]¶
Grid-search competitor parameters against a walk-forward run.
metricdefaults tolog_lossdeliberately. Accuracy is a rank statistic: it only asks which side of 0.5 a prediction landed on, so any parameter that changes confidence without changing order is invisible to it. Pythagorean’s exponent is exactly such a parameter, and tuning it on accuracy reports every value as equally good.- Args:
competitor_class: The rating system to tune. param_grid: Parameter names (without the leading underscore) to sequences of values. periods: Ordered periods, as for
walk_forward(). metric:"log_loss","brier"or"accuracy". **walk_forward_kwargs: Forwarded towalk_forward().- Returns:
Every combination, best first.
- Raises:
ValueError: If
metricis unknown orparam_gridis empty.
- elote.evaluation.walk_forward(competitor_class: Type[BaseCompetitor], periods: Sequence[Sequence[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]], *, competitor_params: Dict[str, Any] | None = None, base_competitor_kwargs: Dict[str, Any] | None = None, comparison_function: Callable[[...], Any] | None = None, score_keys: Tuple[str, str] | None = None, warmup: int = 0, calibration_bins: int = 10) WalkForwardReport[source]¶
Predict each period from everything before it, then learn that period.
Systems that override
BaseCompetitor.apply_rating_period()learn throughLambdaArena.rating_period(); systems that inherit the default implementation keep the existing sequential dataset-training path. A period-native system receives the maximum usable row timestamp asperiod_end(orNonewhen there is none), so time is resolved per period rather than per row. Sequential systems continue to receive each row’s timestamp individually.- Args:
competitor_class: The rating system to evaluate. periods: Ordered periods of dataset rows, as produced by
group_by_period(). competitor_params: Existing class-level knobs to set for the duration of the run,without the leading underscore.
{"default_w2": 100.0}sets_default_w2. Constructor arguments instead belong inbase_competitor_kwargs.base_competitor_kwargs: Constructor keyword arguments for every competitor. comparison_function: Arena comparison function. Defaults to one that reports the
recorded outcome, which is what a dataset row already carries.
- score_keys:
(a_score_key, b_score_key)naming each row’s two point scores, for the margin-aware systems.
- warmup: Leading periods used for fitting but not scored, so a system is not judged
on predictions made with no history.
calibration_bins: Number of equal-width bins for
WalkForwardReport.reliability.- score_keys:
- Returns:
WalkForwardReport: Metrics over every scored, predictable bout.
- Raises:
- ValueError: If
warmupis negative or not smaller than the number of periods, or calibration_binsis not a positive integer.
- ValueError: If