Module pbt_bench
Expand description
The M4.1 PBT oracle’s driver: prove Pbt
discovers a hyperparameter schedule on a nonstationary problem and beats
the best fixed hyperparameter and random search at a matched budget.
This is to PBT what scheduler_bench is to the
pruners (docs/design/09-implementation.md §13): it drives the real study loop
with the real Pbt scheduler, so the Decision::Fork path, the
checkpoint-reference hand-off and the recorded lineage are all exercised
end-to-end, not mocked. PBT only earns its complexity when the best
hyperparameter changes over training (docs/design/04-rl-and-oniro.md §A.2), so
the problem here is deliberately schedule-dependent: no single fixed
learning rate can reach the optimum, a decreasing schedule can, and PBT has to
find that schedule by exploit-and-explore.
§The schedule-dependent problem
Training runs in segments (the resource intervals PBT ranks at). A learning
rate lr used in segment s earns a reward that peaks at a segment-specific
optimum x_opt(s) (in log10(lr) space), and x_opt(s) decreases with
s — a high learning rate is best early, a low one best late
(ScheduleProblem). The objective a run reports is its accumulated reward
across the segments it has trained, so:
- a fixed
lrheld for the whole run can only sit near the optimum of some segments and is off for the rest — its best achievable total is a compromise (ScheduleProblem::best_fixed); - a schedule that tracks
x_opt(s)down earns the per-segment peak at every segment — a total no fixed value reaches (ScheduleProblem::ideal_schedule_total).
§How PBT realizes a schedule through the loop
atune runs the segments as a chain of trials — the run-chaining PBT model
oniro needs (docs/design/04-rl-and-oniro.md §B.2, run_training(steps=interval) →
run_training_resume). Each trial trains one segment, starting from the
carried progress its parent left in a checkpoint, and records a new checkpoint
so an exploit can copy it:
- a root trial (no parent) trains segment 0 with a freshly sampled
lr; - a fork child, created when PBT’s
Decision::Forkfires for an underperformer, copies a top performer’s checkpoint (its carried progress and segment index) and itslrperturbed (Perturb), and trains the next segment.
A lineage root → … → leaf therefore is a learning-rate schedule: the lr
each ancestor used at each segment. PBT’s exploit-and-explore drives the
surviving lineage’s schedule toward the moving optimum, which is exactly what
run_pbt measures and PbtReport records.
§Determinism
Everything is a deterministic function of the study seed: the reward is a pure
function of (segment, lr, seed) (a small, seeded noise term makes it a
genuine — reproducible — function of the seed and exercises
replicate_seed), the
sampler and PBT draw only seeded randomness, and the loop runs single-worker
(parallelism(1)), so the whole population trajectory, the winner and its
lineage replay bit-for-bit (docs/design/09-implementation.md §5).
Structs§
- Member
- One trial’s place in the population, recovered from storage after the run.
- PbtReport
- The result of a PBT run over the schedule problem.
- Schedule
Problem - The nonstationary, schedule-dependent synthetic problem.
Constants§
- LR
- The learning-rate parameter name the objective suggests and the schedule is read from.
Functions§
- random_
search_ best - The best full-schedule total independent random search finds at a budget
matched to a PBT run of
pbt_budgettrials. - run_pbt
- Runs single-worker PBT over
problemforbudgettrials and records the population, the winner and the schedule it discovered. - tuned_
pbt - Builds the tuned PBT scheduler the oracle uses over
ScheduleProblem.