API Reference

Streaming calibration

class streamcal.StreamingIsotonicCalibrator(n_bins=100, decay=None, prior_weight=1.0, *, half_life_seconds=None)[source]

Bounded-memory isotonic calibration for probability streams.

Each update exponentially decays the previous per-bin sufficient statistics, adds the new batch, and computes a weighted isotonic fit. The state therefore has size O(n_bins) regardless of stream length.

Parameters:
reset()[source]

Forget all observations and restore identity calibration.

Return type:

Self

calibrate(probabilities, *, current_time=None)[source]

Calibrate without mutating learned state.

Parameters:
  • probabilities (ArrayLike) – Binary forecast probabilities.

  • current_time (float | None) – Finite seconds on the update timestamp clock. Required in half-life mode and must not precede the latest arrival.

Returns:

Calibrated probabilities in input order.

Return type:

NDArray[np.float64]

update(probabilities, outcomes, *, prediction_times=None, observed_at=None)[source]

Observe labels, weighting late outcomes by their prediction age.

Parameters:
  • probabilities (ArrayLike) – Original forecast probabilities, not calibrated ones.

  • outcomes (ArrayLike) – Corresponding binary labels.

  • prediction_times (ArrayLike | None) – One finite prediction timestamp per observation.

  • observed_at (float | None) – Finite, nondecreasing label arrival timestamp in seconds.

Returns:

This calibrator after the update.

Raises:

ValueError – If timing arguments are missing, misaligned or inconsistent.

Return type:

Self

diagnostics()[source]

Return readiness, data-use, and fixed-state memory diagnostics.

Return type:

CalibratorDiagnostics

class streamcal.CalibratorDiagnostics(is_ready, n_updates, n_observations, effective_observations, occupied_bins, state_bytes)[source]

A snapshot of the state carried by a streaming calibrator.

Parameters:
  • is_ready (bool)

  • n_updates (int)

  • n_observations (int)

  • effective_observations (float)

  • occupied_bins (int)

  • state_bytes (int)

is_ready

Whether at least one outcome batch has been observed.

Type:

bool

n_updates

Number of calls to StreamingIsotonicCalibrator.update().

Type:

int

n_observations

Total number of outcomes observed without decay.

Type:

int

effective_observations

Sum of the decayed bin weights.

Type:

float

occupied_bins

Number of bins with positive effective weight.

Type:

int

state_bytes

Bytes used by the fixed-size numerical state arrays.

Type:

int

Batch references

class streamcal.BatchCalibrator(method='isotonic', *, window_size=None, refit_every=1)[source]

Refit an upstream calibrator on an accumulating or rolling history.

Both-class readiness is required. A one-class window retains the preceding fitted map, or identity if no fit has succeeded.

Parameters:
  • method (str)

  • window_size (int | None)

  • refit_every (int)

reset()[source]

Discard history and the fitted map.

Return type:

Self

freeze()[source]

Keep the fitted map, discard raw history, and ignore future updates.

Return type:

Self

property is_ready: bool

Return whether an upstream fit has succeeded.

property history_bytes: int

Return retained input-array bytes, excluding the fitted model.

property retained_observations: int

Return the number of outcomes available for the next refit.

calibrate(probabilities)[source]

Apply the last completed fit, or identity before readiness.

Parameters:

probabilities (ArrayLike)

Return type:

NDArray[np.float64]

update(probabilities, outcomes)[source]

Retain new observations and refit at the configured cadence.

Parameters:
  • probabilities (ArrayLike)

  • outcomes (ArrayLike)

Return type:

Self

Trade-off evaluation

streamcal.compare_prequential(probabilities, outcomes, configurations, *, batch_size, diagnostic_bins=20, prediction_times=None, label_delays=None, observe_mask=None)[source]

Compare configurations in chronological predict-then-observe order.

Timing fields describe this invocation on this machine. They are useful for local trade-off decisions, not portable performance guarantees.

Parameters:
  • probabilities (ArrayLike) – Raw probabilities in chronological order.

  • outcomes (ArrayLike) – Binary outcomes aligned with probabilities.

  • configurations (Mapping[str, StreamingIsotonicConfig | Callable[[], BaseCalibrator]]) – Names mapped to streaming-isotonic configurations.

  • batch_size (int) – Number of observations predicted before each update.

  • diagnostic_bins (int) – Bin count for the descriptive calibration diagnostic.

  • prediction_times (ArrayLike | None) – Nondecreasing prediction timestamps in seconds. When supplied, River schedules per-event prediction and label delivery; batch_size must be one. Requires the evaluation extra.

  • label_delays (ArrayLike | None) – Nonnegative, finite delay per observation. Defaults to zero.

  • observe_mask (ArrayLike | None) – Boolean vector controlling which labels may update a model. All forecasts are still scored offline; defaults to all True.

Returns:

Quality trajectories and resource measurements for every configuration.

Raises:

ValueError – If inputs, batch size, diagnostic bins, or configuration names are invalid.

Return type:

TradeoffReport

class streamcal.StreamingIsotonicConfig(n_bins=100, decay=None, prior_weight=1.0, half_life_seconds=None)[source]

Configuration used by compare_prequential().

Parameters:
build()[source]

Construct and validate a fresh calibrator.

Return type:

StreamingIsotonicCalibrator

class streamcal.TradeoffReport(n_observations, batch_size, results)[source]

Prequential results for a configuration comparison.

Parameters:
pareto_frontier()[source]

Return nondominated Brier, serialized-size, and update-time choices.

Return type:

tuple[ConfigurationResult, …]

select(*, objective='brier', max_brier_degradation=None, max_serialized_bytes=None, max_update_ns_per_observation=None, max_predict_ms=None, max_update_ms=None)[source]

Minimize the chosen error subject to validation-measured budgets.

Parameters:
  • objective (str) – brier, log_loss, or binned_calibration_error.

  • max_brier_degradation (float | None) – Optional increase over the best Brier in this report.

  • max_serialized_bytes (int | None) – Optional nonnegative integer state budget in bytes.

  • max_update_ns_per_observation (float | None) – Optional measured update-time limit.

  • max_predict_ms (float | None) – Maximum p95 prediction-call latency in milliseconds.

  • max_update_ms (float | None) – Maximum p95 update-call latency in milliseconds.

Returns:

Lowest-error feasible result, breaking ties by size then name, or None.

Raises:
  • ValueError – If the objective is unknown or a budget is negative.

  • TypeError – If the serialized-state budget is not an integer.

Return type:

ConfigurationResult | None

paired_interval(candidate, reference, *, block_size, replicates=2000, seed=2026)[source]

Return paired mean Brier difference and circular-block 95% interval.

Requires the evaluation extra. Block size is in observations in prediction order. Choose it for the dependence scale of your stream.

Parameters:
  • candidate (str) – Name of the candidate method.

  • reference (str) – Name of the reference method.

  • block_size (int) – Number of consecutive observations per bootstrap block.

  • replicates (int) – Number of bootstrap replicates.

  • seed (int) – Bootstrap random seed.

Returns:

Mean difference, lower endpoint, and upper endpoint.

Raises:

ValueError – If observations are absent or the block exceeds their count.

Return type:

tuple[float, float, float]

class streamcal.ConfigurationResult(name, brier, log_loss, binned_calibration_error, roc_auc, predict_ns_per_observation, update_ns_per_observation, state_bytes, batch_brier, cumulative_brier, serialized_bytes=0, squared_errors=(), p95_predict_ms=0.0, p95_update_ms=0.0)[source]

Quality and resource measurements for one configuration.

Parameters:

Diagnostics

streamcal.brier_score(outcomes, probabilities)[source]

Return mean squared probability error for binary outcomes.

Parameters:
  • outcomes (ArrayLike) – Observed binary outcomes.

  • probabilities (ArrayLike) – Forecast probabilities aligned with outcomes.

Returns:

Mean squared probability error.

Return type:

float

streamcal.binned_calibration_error(outcomes, probabilities, n_bins=20)[source]

Return equal-width binned absolute calibration error.

This diagnostic depends on both sample size and n_bins. It is useful for visualization and comparison at a fixed design, but it is not a proper score and should not be used alone to select a probabilistic forecast.

Parameters:
  • outcomes (ArrayLike) – Observed binary outcomes.

  • probabilities (ArrayLike) – Forecast probabilities aligned with outcomes.

  • n_bins (int) – Number of equal-width bins over [0, 1].

Returns:

Observation-weighted absolute gap between bin means.

Return type:

float