Skip to main content

Module pbt_bench

Module pbt_bench 

Expand description

The M4.1 PBT oracle’s driver: prove Pbt discovers a hyperparameter schedule on a nonstationary problem and beats the best fixed hyperparameter and random search at a matched budget.

This is to PBT what scheduler_bench is to the pruners (docs/design/09-implementation.md §13): it drives the real study loop with the real Pbt scheduler, so the Decision::Fork path, the checkpoint-reference hand-off and the recorded lineage are all exercised end-to-end, not mocked. PBT only earns its complexity when the best hyperparameter changes over training (docs/design/04-rl-and-oniro.md §A.2), so the problem here is deliberately schedule-dependent: no single fixed learning rate can reach the optimum, a decreasing schedule can, and PBT has to find that schedule by exploit-and-explore.

§The schedule-dependent problem

Training runs in segments (the resource intervals PBT ranks at). A learning rate lr used in segment s earns a reward that peaks at a segment-specific optimum x_opt(s) (in log10(lr) space), and x_opt(s) decreases with s — a high learning rate is best early, a low one best late (ScheduleProblem). The objective a run reports is its accumulated reward across the segments it has trained, so:

  • a fixed lr held for the whole run can only sit near the optimum of some segments and is off for the rest — its best achievable total is a compromise (ScheduleProblem::best_fixed);
  • a schedule that tracks x_opt(s) down earns the per-segment peak at every segment — a total no fixed value reaches (ScheduleProblem::ideal_schedule_total).

§How PBT realizes a schedule through the loop

atune runs the segments as a chain of trials — the run-chaining PBT model oniro needs (docs/design/04-rl-and-oniro.md §B.2, run_training(steps=interval) → run_training_resume). Each trial trains one segment, starting from the carried progress its parent left in a checkpoint, and records a new checkpoint so an exploit can copy it:

  • a root trial (no parent) trains segment 0 with a freshly sampled lr;
  • a fork child, created when PBT’s Decision::Fork fires for an underperformer, copies a top performer’s checkpoint (its carried progress and segment index) and its lr perturbed (Perturb), and trains the next segment.

A lineage root → … → leaf therefore is a learning-rate schedule: the lr each ancestor used at each segment. PBT’s exploit-and-explore drives the surviving lineage’s schedule toward the moving optimum, which is exactly what run_pbt measures and PbtReport records.

§Determinism

Everything is a deterministic function of the study seed: the reward is a pure function of (segment, lr, seed) (a small, seeded noise term makes it a genuine — reproducible — function of the seed and exercises replicate_seed), the sampler and PBT draw only seeded randomness, and the loop runs single-worker (parallelism(1)), so the whole population trajectory, the winner and its lineage replay bit-for-bit (docs/design/09-implementation.md §5).

Structs§

Member
One trial’s place in the population, recovered from storage after the run.
PbtReport
The result of a PBT run over the schedule problem.
ScheduleProblem
The nonstationary, schedule-dependent synthetic problem.

Constants§

LR
The learning-rate parameter name the objective suggests and the schedule is read from.

Functions§

random_search_best
The best full-schedule total independent random search finds at a budget matched to a PBT run of pbt_budget trials.
run_pbt
Runs single-worker PBT over problem for budget trials and records the population, the winner and the schedule it discovered.
tuned_pbt
Builds the tuned PBT scheduler the oracle uses over ScheduleProblem.