Write a scheduler¶
A scheduler answers a different question from a sampler, at a different moment: what should happen to this trial, now? It is handed every intermediate report a running trial makes, and it replies with one of four decisions.
The thing to know before you start is that a scheduler is a superset of a
pruner. Conflating the two makes pause, resume and fork inexpressible, which
rules out the whole population-based family; here a pruner is simply a scheduler
that only ever answers continue or prune. Pbt and FreezeThaw implement the
same trait you are about to implement, and they use it to fork and to revive
trials rather than only to stop them.
The four decisions¶
| Decision | What the study loop does | Stops the trial |
|---|---|---|
Continue |
Nothing. The trial runs on. | no |
Prune |
The trial stops now and is recorded as pruned, keeping its last intermediate value as its objective value. | yes |
Pause |
The trial transitions to a resumable, non-terminal Paused state that keeps its checkpoint reference. It may be woken later. |
yes |
Fork |
A new trial is created from the named parent's parameters and checkpoint, with a mutation applied. The trial that reported is not stopped. | no |
Two of those carry obligations that are easy to miss.
Fork may only mutate scalars. The child starts from the parent's checkpoint
reference, and a perturbed shape or categorical parameter would make that
checkpoint unloadable — so a non-scalar mutation is refused loudly rather than
producing a child that cannot resume.
Pause needs a way back. A pause is non-terminal, so it never triggers the
end-of-trial hook; a scheduler that pauses trials nothing later revives strands
them, and the paused pool is never drained. That is what resume_candidates
below exists for, and FreezeThaw is exactly that regime — a slot is freed by a
pause, not by a finish.
What the trait asks for¶
Seven methods, of which one is required:
| Method | Required | What it does |
|---|---|---|
on_report |
yes | The decision. Receives the study view, the reporting trial, the step, and a slice of values — pruning is multi-objective-aware from the start, and a scheduler is entitled to look at every objective. |
on_trial_end |
no | Called when a trial reaches a terminal state; returns commands. Where rung tables are updated and where a promotion decides to wake a paused trial. |
resume_candidates |
no | Paused trials worth waking on a freed work slot, in priority order. Consulted before the loop creates fresh work. |
state |
no | Persistable state as a typed, versioned blob, or None. |
restore_state |
no | Called once, before any decision, when a study is created or resumed. |
scripted_capability |
no | What the scripted ask/tell path may do with this scheduler; the default declines it. |
fan_report_mode |
no | How multi-seed fans report to on_report; the default rejects a fan before trial creation (see below). |
step is counted in the study's declared resource unit, so multi-fidelity
arithmetic is explicit rather than "whatever step meant to the caller".
Three constraints on on_report follow from where it runs — inside the
objective's inner loop, on every worker:
- It must be cheap. Every report pays for it.
- It must not go back to storage beyond the view it was handed.
- An error is not a trial failure. The loop reads an error from a scheduler as
"no decision" and lets the trial continue. Return
Error::Schedulerrather than panicking, and never rely on an error to stop a trial.
What you can and cannot see¶
The snapshot does not contain the trial that is reporting. It was taken when
that trial was created. So a scheduler that needs the reporting trial's own
history — a patience counter, a paired test, a curve fit — keeps that bookkeeping
itself. The built-ins that need it (Patient, WilcoxonPruner, FreezeThaw) do
exactly that, behind interior mutability, because the trait hands you &self.
A multi-seed fan reports only aligned aggregates. Per-replicate reports are
captured rather than streamed. After all required replicates finish, equal step
sets and arities are verified and on_report receives one deterministic
aggregate per step; its decision applies to the whole logical trial. Opt in with
fan_report_mode() == FanReportMode::AlignedAggregate; the default rejects a
fan before trial creation. When the trial ends, on_trial_end receives it with
its per-seed fan filled in — but the other trials reachable through the study
view carry an empty fan, deliberately. A comparison against an incumbent's fan
therefore remains caller-side.
Registering it¶
One line, and it is the same line a built-in uses. The builder takes an
Arc<dyn Scheduler>:
let study = Study::builder()
.parallelism(THREADS)
.budget(Budget::trials(TRIALS))
// Rank every INTERVAL units; Jaderberg defaults for the rest (the bottom
// 20% exploit a random member of the top 20%, perturbing by 0.8 / 1.2).
.scheduler(Arc::new(Pbt::new(u64::from(INTERVAL))?))
.create(StudyConfig::new(STUDY).with_seed(SEED))?;
study.optimize(train)?;
// The genealogy, rebuilt from storage: who forked from whom, and with what
// perturbation. PBT persists no scheduler state — the tree *is* the record.
let lineage = study.lineage()?;
That block is Pbt, a fork-capable scheduler, shown Rust-only for the same reason
Write a sampler is: the Python bindings hand out scheduler
handles built in Rust, so the trait is implemented in Rust and then used from
either language.
If you keep state¶
Rung tables, populations, promotion ladders: return them from state as a typed,
versioned blob and take them back in restore_state. The framework stores it
under the study's scheduler scope and hands it back on resume, so a study
continues where the previous handle left off instead of restarting its
bookkeeping.
Not everything needs it. Every built-in pruner is stateless by recomputation — it
rebuilds what it needs from the study view — and so persists nothing, which is why
the default implementations of both methods are inert. Pbt keeps no scheduler
state either: the fork tree in storage is the record. FreezeThaw is the one
built-in that does persist, because the partial learning curves it fits cannot be
recovered from the view. Reach for a blob when a decision genuinely cannot be
recomputed from history, not by reflex.
Determinism, honestly¶
Every real scheduler decides from the trials it can see, and which trials those are depends on completion order — which under parallelism is thread scheduling. A scheduled study is therefore replayable (its recorded history explains every prune, pause and fork) rather than pre-determined (computable from the seed before it runs). Under a single worker with a fixed seed it is fully reproducible.
Say which of those your scheduler offers, in its own documentation. Every built-in does, and Determinism is the contract those words refer to.
One trap that is not yours to fix¶
A pruned trial's objective value is its last intermediate report. If a study's reported metric gets worse as the resource grows, the trials your scheduler stopped early will beat the ones that ran to completion, and any sampler learning from that history is steered away from the good region.
That is worth knowing while writing a scheduler so that you do not try to correct it in the scheduler — the fix belongs in the objective, and Prune and schedule states it in full, with the measurement that made the point.
Where to go next¶
| If you want to | Go to |
|---|---|
| See the seam from the user's side | Prune and schedule |
| See fork and revive in a real study | Population-based training |
| Understand how the two seams relate | Samplers and schedulers |
| Choose a trial's parameters instead | Write a sampler |
| Ship it as a crate someone can depend on | Publish a plugin crate |
| Read the trait, method by method | Rust API |