Skip to main content

Module online_bench

Module online_bench 

Expand description

The online-tuner ablation’s driver: a synthetic non-stationary in-loop tuning problem, and the three strategies compared on it (M6.0, docs/design/09-implementation.md §15, §8.3; docs/design/04-rl-and-oniro.md §A.3).

The whole point of an in-loop tuner is a reward whose optimum moves during training — a single fixed hyperparameter cannot be right the whole way. So the problem here is a knob x ∈ [0, 1] whose ideal value drifts across three equal regimes, and the reward is how close the knob sits to that moving target:

x*(round) = ARMS[(round / REGIME_LEN) mod 3]      // the drifting optimum
reward    = 1 − (x − x*(round))²  + noise·ξ(seed, round)

Three strategies are driven over the same reward realizations (same seed ⇒ same noise), so the comparison is fair:

  • online — an OnlineTuner with the clustered sliding-window UCB Mabc policy over the discrete ARMS; it should track the drift.
  • static — the single best fixed arm in hindsight (the strongest static baseline: it already knows which one arm maximizes total reward).
  • pbt-lite — an OnlineTuner with the PbtLite continuous-knob policy, chasing the drift by perturbation.

Everything is a pure function of (seed, noise); on the noise-free instance it is a pure function of nothing, so the mechanical claims are seed-free (docs/design/09-implementation.md §5). The online_oracle integration test turns these reports into the §8.3 ablation.

Structs§

AblationRow
One seed’s ablation row: the total reward each strategy achieved.
OnlineReport
What one online run produced.

Constants§

ARMS
The three candidate knob values — also the three regime optima, so the best arm rotates 0 → 1 → 2 across the run.
N_REGIMES
The number of regimes (and distinct optima) in one run.
N_ROUNDS
The total number of rounds in one run.
REGIME_LEN
The number of rounds per regime.

Functions§

ablation_row
Measures all three strategies on one seed with a shared configuration.
optimal_arm
The regime (and therefore the optimal arm index) active at round.
reward
The reward for placing the knob at x on round, under noise noise_amp keyed by seed. Noise-free when noise_amp == 0 (then seed is ignored).
run_online
Runs the online UCB tuner over the drift problem and reports its behaviour.
run_pbt_lite
Runs the PBT-lite continuous-knob tuner over the drift problem; returns its total reward. The knob’s feasible range is [0, 1] and it starts mid-range.
run_static_best
The best fixed arm in hindsight and its total reward.
x_star
The drifting optimum x*(round).