Run population-based training¶
Goal: let good configurations copy themselves mid-run, instead of running every configuration to completion in isolation.
Ordinary tuning treats a trial as atomic: pick a configuration, run it, score it, pick the next. Population-based training breaks that. Trials run concurrently as a population; periodically each is ranked against the others, and an underperformer exploits a winner — copying its hyperparameters and its checkpoint — then explores by perturbing what it copied.
The consequence worth understanding before you use it: the answer PBT gives you is usually a fork of a fork, and the hyperparameters that produced it were never a single fixed configuration. They are a schedule, discovered while training.
Attach the scheduler¶
let study = Study::builder()
.parallelism(THREADS)
.budget(Budget::trials(TRIALS))
// Rank every INTERVAL units; Jaderberg defaults for the rest (the bottom
// 20% exploit a random member of the top 20%, perturbing by 0.8 / 1.2).
.scheduler(Arc::new(Pbt::new(u64::from(INTERVAL))?))
.create(StudyConfig::new(STUDY).with_seed(SEED))?;
study.optimize(train)?;
// The genealogy, rebuilt from storage: who forked from whom, and with what
// perturbation. PBT persists no scheduler state — the tree *is* the record.
let lineage = study.lineage()?;
study = atune.create_study(
direction="minimize",
# Rank every INTERVAL units; Jaderberg defaults for the rest (the bottom
# 20% exploit a random member of the top 20%, perturbing by 0.8 / 1.2).
scheduler=atune.schedulers.Pbt(INTERVAL),
seed=SEED,
name=STUDY,
)
study.optimize(objective, n_trials=TRIALS, n_jobs=JOBS)
# The genealogy, rebuilt from storage: who forked from whom, and with what
# perturbation. PBT persists no scheduler state — the tree *is* the record.
lineage = study.lineage()
Pbt::new(interval) ranks the population every interval resource units. The
rest are Jaderberg's defaults: the bottom fifth exploits a random member of the
top fifth, and perturbation multiplies a copied scalar by 0.8 or 1.2.
Nothing else about the study changes. PBT is a Scheduler, so it attaches exactly
where a pruner would — the difference is what it may return. A pruner can say
stop; PBT can also say fork, and the study loop then materialises a new trial
whose parameters and checkpoint come from another one.
The objective has to be resumable¶
This is the requirement that makes PBT different to configure, and it is not optional. Exploit copies a trial's state, not just its numbers. If your objective cannot start from another trial's checkpoint, exploiting a winner gives you its learning rate and none of its progress — which is strictly worse than leaving the trial alone, because you have discarded the work it had already done.
So an objective under PBT does three things a plain objective does not:
/// The loss this learning rate converges to — what it is *capable* of.
///
/// A quadratic bowl around [`LR_STAR`]: exactly [`FLOOR_MIN`] at the optimum and
/// worse either side of it. Pure arithmetic, in the same order as the Python arm,
/// so both arms compute the same bits.
fn floor_for(lr: f64) -> f64 {
let off = lr - LR_STAR;
FLOOR_MIN + FLOOR_SLOPE * (off * off)
}
/// Trains one segment, warm-starting from the parent's checkpoint if there is one.
fn train(ctx: &mut TrialCtx<'_>) -> Result<Outcome> {
let lr = ctx.suggest_f64("lr", LR_LOW..=LR_HIGH, Scale::Linear)?;
// The exploit half of PBT: a fork child inherits its parent's checkpoint
// reference. atune stored the string, not the state — here the state is one
// number, so the string is `{loss:?}` and `parse` recovers its exact bits.
// Parsed into an owned `f64` immediately: the borrow of the context ends
// here, before the loop needs it mutably again.
let mut loss = match ctx.parent_checkpoint() {
None => INITIAL_LOSS,
Some(reference) => reference.parse::<f64>().map_err(|error| {
Error::Objective(
format!("checkpoint reference {reference:?} is not a loss: {error}").into(),
)
})?,
};
let floor = floor_for(lr);
for step in 1..=STEPS {
loss = floor + (loss - floor) * DECAY;
if step % INTERVAL == 0 {
// Checkpoint *before* reporting: `report` is the scheduler's decision
// point, and a fork decided there copies the reference recorded here.
ctx.record_checkpoint(u64::from(step), format!("{loss:?}"))?;
ctx.report(u64::from(step), &[loss])?;
}
}
Ok(loss.into())
}
def floor_for(lr: float) -> float:
"""The loss this learning rate converges to — what it is *capable* of.
A quadratic bowl around ``LR_STAR``: exactly ``FLOOR_MIN`` at the optimum and
worse either side of it. Pure arithmetic, in the same order as the Rust arm,
so both arms compute the same bits.
"""
off = lr - LR_STAR
return FLOOR_MIN + FLOOR_SLOPE * (off * off)
def objective(trial: atune.Trial) -> float:
"""Trains one segment, warm-starting from the parent's checkpoint if any."""
lr = trial.suggest_float("lr", LR_LOW, LR_HIGH)
# The exploit half of PBT: a fork child inherits its parent's checkpoint
# reference. atune stored the string, not the state — here the state is one
# number, so the string is `repr(loss)` and `float` recovers its exact bits.
reference = trial.parent_checkpoint
loss = INITIAL_LOSS if reference is None else float(reference)
floor = floor_for(lr)
for step in range(1, STEPS + 1):
loss = floor + (loss - floor) * DECAY
if step % INTERVAL == 0:
# Checkpoint *before* reporting: `report` is the scheduler's decision
# point, and a fork decided there copies the reference recorded here.
trial.record_checkpoint(step, repr(loss))
trial.report(loss, step)
return loss
It looks for an inherited checkpoint and warm-starts from it; it reports at intervals, so the scheduler has something to rank; and it writes a checkpoint the next generation can inherit. In a real trainer that reference is a path to model weights. In the example it is deliberately small enough to read, so the mechanism is visible rather than buried in a framework.
Read the genealogy¶
A flat list of trials is the wrong shape for a PBT study — it hides the only
structure that matters. study.lineage() rebuilds the tree from storage: who
forked from whom, and — on the Rust side — with which perturbation; the Python
view carries each node's parent, fork flag and value.
PBT persists no scheduler state of its own. The tree is the record, which is also why a PBT study resumes correctly: reopen the storage and the population's history is all there.
Make the population actually pay¶
Two ways a PBT study can run, report success, and demonstrate nothing.
The segment must be short enough that a warm start still matters. If every trial reaches its floor before the first ranking, exploiting a winner inherits a finished job, and a fork becomes indistinguishable from a fresh trial with the same parameters. Measured on the example above: at eight steps with decay 0.6, about 1.7% of the initial gap survives to the ranking — which is exactly what makes a fork measurably better than a cold trial that drew the same learning rate. Longer segments erase the effect, not because PBT stopped working but because there was nothing left to inherit.
The population must be large enough to have a top and a bottom fifth. The
population at a decision is every trial that has reported at that step and is
still rankable — running and paused trials and finished ones alike — so it
accumulates over the study rather than being capped by parallelism (the
example runs single-threaded and still forks). What matters is that enough
trials report at the same aligned steps: with only a handful having reached
a step, "the bottom 20%" is less than one trial and no exploit fires there
yet.
PB2, and when to reach for it¶
Pb2 replaces PBT's random perturbation with one chosen by a time-varying
GP-bandit: instead of multiplying by 0.8 or 1.2, it proposes the perturbation its
model expects to help most given what the population has done so far. It is behind
the pb2 feature, which implies gp — see the
feature reference.
Reach for it when perturbation is the bottleneck: a population that exploits correctly but explores badly. It is not a drop-in improvement, because it adds a model whose own assumptions can be wrong.
What this costs you¶
- The result is a schedule, not a configuration. You cannot hand someone "the best hyperparameters" from a PBT study and expect them to reproduce your number — they need the trajectory. Report the lineage alongside the winner.
- Forking is history-dependent. Which trial sits in the bottom fifth at a ranking depends on what has finished, so a PBT study with several workers is not reproducible trial-for-trial. Determinism states what does still hold.
- Checkpoints cost storage. A population that forks every interval writes a checkpoint per trial per interval. That is a real disk budget, and the reason checkpoint references are references rather than payloads.
Next¶
- Prune and schedule trials — the same seam, stopping rather than forking.
- Samplers and schedulers — what a scheduler is allowed to decide.
- Storage — where checkpoints and lineage live.