Evaluation API Reference

Walk-forward evaluation

benchmark_competitors compares systems on one frozen train/test split. Walk-forward evaluation predicts each period from everything before it and then learns that period, and it reports log loss and Brier score alongside accuracy. compare_walk_forward runs that protocol for several systems over the same periods and warmup, so the numbers are comparable.

from elote import (
    EloCompetitor, GlickoBoostCompetitor, SyntheticDataset,
    compare_walk_forward, group_by_period,
)

periods = group_by_period(SyntheticDataset(num_competitors=20, num_matchups=1000, seed=1).load())
result = compare_walk_forward(
    {
        "elo": EloCompetitor,
        "elo_k32": (EloCompetitor, {"competitor_params": {"k_factor": 32}}),
        "glicko_boost": GlickoBoostCompetitor,
    },
    periods,
    warmup=4,
)
print(result)
result.ranking()[0]["system"]  # best log loss

Each system is a class, or a (class, options) pair where options may hold competitor_params and base_competitor_kwargs. Invalid systems raise before anything is evaluated. Walk-forward numbers are not interchangeable with frozen-split numbers: they score a different population of bouts, and warmup decides how much early history is excluded. same_population is False when systems scored different bout counts.

Walk-forward evaluation and hyperparameter search for rating systems.

evaluate_competitor() trains on one split and then predicts a held-out split with frozen ratings. That answers “how well do these ratings survive going stale”, which is a real question but rarely the one being asked. The usual question is how a system performs in the way it would actually be used: predict the next round of results from everything that has happened so far, then fold those results in and step forward.

This module provides that protocol, the metrics that can see a system’s calibration as well as its picks, and a grid search over competitor parameters.

Example:
>>> from elote import EloCompetitor, walk_forward, group_by_period
>>> periods = group_by_period(rows)
>>> report = walk_forward(EloCompetitor, periods)
>>> report.accuracy, report.log_loss
class elote.evaluation.ReliabilityBin(lower: float, upper: float, count: int, mean_predicted: float | None, observed_rate: float | None)[source]

Bases: object

One equal-width probability bin of a reliability table.

The bin covers [lower, upper); the final bin of a table also includes 1.0.

Attributes:

lower: Inclusive lower bound. upper: Exclusive upper bound (inclusive for the final bin). count: Scored predictions whose probability fell in the bin. mean_predicted: Mean predicted probability that the first side wins, or None if empty. observed_rate: Fraction of the bin’s bouts the first side actually won, or None if empty.

count: int
lower: float
mean_predicted: float | None
observed_rate: float | None
upper: float
class elote.evaluation.TuningResult(params: Dict[str, Any], report: WalkForwardReport)[source]

Bases: object

One point of a tune() grid search.

params: Dict[str, Any]
report: WalkForwardReport
class elote.evaluation.WalkForwardComparison(reports: Dict[str, WalkForwardReport], warmup: int, periods: int, same_population: bool)[source]

Bases: object

Walk-forward reports for several systems run over the same periods and warmup.

Attributes:

reports: Label to WalkForwardReport, in the order the systems were given. warmup: Leading periods used for fitting but not scored, shared by every system. periods: Number of periods every system was run over. same_population: True when every system scored the same number of bouts and

skipped and drew the same numbers. When False the figures describe different row populations and should not be compared as-is.

periods: int
ranking() → List[Dict[str, Any]][source]

Rows sorted by log loss, best first; systems with no scored bouts (NaN) come last.

Every row carries the protocol (warmup, periods) and the scored-row counts.

reports: Dict[str, WalkForwardReport]
same_population: bool
warmup: int
class elote.evaluation.WalkForwardReport(predictions: int, skipped: int, draws: int, accuracy: float, log_loss: float, brier: float, by_period: Tuple[Tuple[int, int, float], ...] = (), reliability: Tuple[ReliabilityBin, ...] = ())[source]

Bases: object

Metrics from a walk-forward run.

Attributes:

predictions: Bouts that were both scored and predictable. skipped: Bouts skipped because a competitor had not been seen yet. draws: Drawn bouts, excluded from every metric below. accuracy: Fraction of predictions on the correct side of 0.5. log_loss: Mean negative log likelihood. Sees calibration; accuracy does not. brier: Mean squared error of the predicted probability. by_period: (period_index, predictions, accuracy) per scored period. reliability: Equal-width ReliabilityBin records over exactly the

predictions population (decisive, predictable, post-warmup bouts; draws are excluded). Binned on the original prediction, before the log-loss clamp.

accuracy: float
brier: float
by_period: Tuple[Tuple[int, int, float], ...] = ()
draws: int
log_loss: float
predictions: int
reliability: Tuple[ReliabilityBin, ...] = ()
skipped: int
elote.evaluation.compare_walk_forward(systems: Dict[str, Any], periods: Sequence[Sequence[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]], *, warmup: int = 0, score_keys: Tuple[str, str] | None = None) → WalkForwardComparison[source]

Run walk_forward() for each system on the same periods and warmup.

Args:
systems: Label to a competitor class, or to a (class, options) pair where

options may hold competitor_params and base_competitor_kwargs exactly as walk_forward() takes them.

periods: Ordered periods of dataset rows, as produced by group_by_period(). warmup: Leading periods used for fitting but not scored, applied to every system. score_keys: (a_score_key, b_score_key) naming each row’s two point scores.

Returns:

WalkForwardComparison: Each system’s report, a log-loss ranking and a printable table.

Raises:
InvalidParameterException: If systems is empty or any entry is malformed, is not a

competitor class, or names a parameter its class does not have. Raised before any system is evaluated.

ValueError: If warmup is negative or not smaller than the number of periods.

elote.evaluation.group_by_period(rows: Iterable[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]], key: Callable[[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]], Any] | None = None) → List[List[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]][source]

Group dataset rows into chronologically ordered periods.

A period is the unit of “predict, then learn”: everything inside one is predicted before any of it is used for fitting, which is what stops a result informing a bet placed on the same afternoon.

Args:

rows: Dataset rows, in any order. key: Maps a row to its period. Defaults to the ISO calendar week of the row’s

timestamp, which suits weekly league sports. Rows without a usable timestamp are collected into one leading period.

Returns:

A list of periods, each a list of rows, ordered by period key.

elote.evaluation.tune(competitor_class: Type[BaseCompetitor], param_grid: Dict[str, Sequence[Any]], periods: Sequence[Sequence[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]], *, metric: str = 'log_loss', **walk_forward_kwargs: Any) → List[TuningResult][source]

Grid-search competitor parameters against a walk-forward run.

metric defaults to log_loss deliberately. Accuracy is a rank statistic: it only asks which side of 0.5 a prediction landed on, so any parameter that changes confidence without changing order is invisible to it. Pythagorean’s exponent is exactly such a parameter, and tuning it on accuracy reports every value as equally good.

Args:

competitor_class: The rating system to tune. param_grid: Parameter names (without the leading underscore) to sequences of values. periods: Ordered periods, as for walk_forward(). metric: "log_loss", "brier" or "accuracy". **walk_forward_kwargs: Forwarded to walk_forward().

Returns:

Every combination, best first.

Raises:

ValueError: If metric is unknown or param_grid is empty.

elote.evaluation.walk_forward(competitor_class: Type[BaseCompetitor], periods: Sequence[Sequence[Tuple[Any, Any, float, datetime | None, Dict[str, Any] | None]]], *, competitor_params: Dict[str, Any] | None = None, base_competitor_kwargs: Dict[str, Any] | None = None, comparison_function: Callable[[...], Any] | None = None, score_keys: Tuple[str, str] | None = None, warmup: int = 0, calibration_bins: int = 10) → WalkForwardReport[source]

Predict each period from everything before it, then learn that period.

Systems that override BaseCompetitor.apply_rating_period() learn through LambdaArena.rating_period(); systems that inherit the default implementation keep the existing sequential dataset-training path. A period-native system receives the maximum usable row timestamp as period_end (or None when there is none), so time is resolved per period rather than per row. Sequential systems continue to receive each row’s timestamp individually.

Args:

competitor_class: The rating system to evaluate. periods: Ordered periods of dataset rows, as produced by group_by_period(). competitor_params: Existing class-level knobs to set for the duration of the run,

without the leading underscore. {"default_w2": 100.0} sets _default_w2. Constructor arguments instead belong in base_competitor_kwargs.

base_competitor_kwargs: Constructor keyword arguments for every competitor. comparison_function: Arena comparison function. Defaults to one that reports the

recorded outcome, which is what a dataset row already carries.

score_keys: (a_score_key, b_score_key) naming each row’s two point scores, for

the margin-aware systems.

warmup: Leading periods used for fitting but not scored, so a system is not judged

on predictions made with no history.

calibration_bins: Number of equal-width bins for WalkForwardReport.reliability.

Returns:

WalkForwardReport: Metrics over every scored, predictable bout.

Raises:
ValueError: If warmup is negative or not smaller than the number of periods, or

calibration_bins is not a positive integer.