Skip to content

Write a scheduler

A scheduler answers a different question from a sampler, at a different moment: what should happen to this trial, now? It is handed every intermediate report a running trial makes, and it replies with one of four decisions.

The thing to know before you start is that a scheduler is a superset of a pruner. Conflating the two makes pause, resume and fork inexpressible, which rules out the whole population-based family; here a pruner is simply a scheduler that only ever answers continue or prune. Pbt and FreezeThaw implement the same trait you are about to implement, and they use it to fork and to revive trials rather than only to stop them.

The four decisions

Decision What the study loop does Stops the trial
Continue Nothing. The trial runs on. no
Prune The trial stops now and is recorded as pruned, keeping its last intermediate value as its objective value. yes
Pause The trial transitions to a resumable, non-terminal Paused state that keeps its checkpoint reference. It may be woken later. yes
Fork A new trial is created from the named parent's parameters and checkpoint, with a mutation applied. The trial that reported is not stopped. no

Two of those carry obligations that are easy to miss.

Fork may only mutate scalars. The child starts from the parent's checkpoint reference, and a perturbed shape or categorical parameter would make that checkpoint unloadable — so a non-scalar mutation is refused loudly rather than producing a child that cannot resume.

Pause needs a way back. A pause is non-terminal, so it never triggers the end-of-trial hook; a scheduler that pauses trials nothing later revives strands them, and the paused pool is never drained. That is what resume_candidates below exists for, and FreezeThaw is exactly that regime — a slot is freed by a pause, not by a finish.

What the trait asks for

Seven methods, of which one is required:

Method Required What it does
on_report yes The decision. Receives the study view, the reporting trial, the step, and a slice of values — pruning is multi-objective-aware from the start, and a scheduler is entitled to look at every objective.
on_trial_end no Called when a trial reaches a terminal state; returns commands. Where rung tables are updated and where a promotion decides to wake a paused trial.
resume_candidates no Paused trials worth waking on a freed work slot, in priority order. Consulted before the loop creates fresh work.
state no Persistable state as a typed, versioned blob, or None.
restore_state no Called once, before any decision, when a study is created or resumed.
scripted_capability no What the scripted ask/tell path may do with this scheduler; the default declines it.
fan_report_mode no How multi-seed fans report to on_report; the default rejects a fan before trial creation (see below).

step is counted in the study's declared resource unit, so multi-fidelity arithmetic is explicit rather than "whatever step meant to the caller".

Three constraints on on_report follow from where it runs — inside the objective's inner loop, on every worker:

  • It must be cheap. Every report pays for it.
  • It must not go back to storage beyond the view it was handed.
  • An error is not a trial failure. The loop reads an error from a scheduler as "no decision" and lets the trial continue. Return Error::Scheduler rather than panicking, and never rely on an error to stop a trial.

What you can and cannot see

The snapshot does not contain the trial that is reporting. It was taken when that trial was created. So a scheduler that needs the reporting trial's own history — a patience counter, a paired test, a curve fit — keeps that bookkeeping itself. The built-ins that need it (Patient, WilcoxonPruner, FreezeThaw) do exactly that, behind interior mutability, because the trait hands you &self.

A multi-seed fan reports only aligned aggregates. Per-replicate reports are captured rather than streamed. After all required replicates finish, equal step sets and arities are verified and on_report receives one deterministic aggregate per step; its decision applies to the whole logical trial. Opt in with fan_report_mode() == FanReportMode::AlignedAggregate; the default rejects a fan before trial creation. When the trial ends, on_trial_end receives it with its per-seed fan filled in — but the other trials reachable through the study view carry an empty fan, deliberately. A comparison against an incumbent's fan therefore remains caller-side.

Registering it

One line, and it is the same line a built-in uses. The builder takes an Arc<dyn Scheduler>:

let study = Study::builder()
    .parallelism(THREADS)
    .budget(Budget::trials(TRIALS))
    // Rank every INTERVAL units; Jaderberg defaults for the rest (the bottom
    // 20% exploit a random member of the top 20%, perturbing by 0.8 / 1.2).
    .scheduler(Arc::new(Pbt::new(u64::from(INTERVAL))?))
    .create(StudyConfig::new(STUDY).with_seed(SEED))?;

study.optimize(train)?;

// The genealogy, rebuilt from storage: who forked from whom, and with what
// perturbation. PBT persists no scheduler state — the tree *is* the record.
let lineage = study.lineage()?;

That block is Pbt, a fork-capable scheduler, shown Rust-only for the same reason Write a sampler is: the Python bindings hand out scheduler handles built in Rust, so the trait is implemented in Rust and then used from either language.

If you keep state

Rung tables, populations, promotion ladders: return them from state as a typed, versioned blob and take them back in restore_state. The framework stores it under the study's scheduler scope and hands it back on resume, so a study continues where the previous handle left off instead of restarting its bookkeeping.

Not everything needs it. Every built-in pruner is stateless by recomputation — it rebuilds what it needs from the study view — and so persists nothing, which is why the default implementations of both methods are inert. Pbt keeps no scheduler state either: the fork tree in storage is the record. FreezeThaw is the one built-in that does persist, because the partial learning curves it fits cannot be recovered from the view. Reach for a blob when a decision genuinely cannot be recomputed from history, not by reflex.

Determinism, honestly

Every real scheduler decides from the trials it can see, and which trials those are depends on completion order — which under parallelism is thread scheduling. A scheduled study is therefore replayable (its recorded history explains every prune, pause and fork) rather than pre-determined (computable from the seed before it runs). Under a single worker with a fixed seed it is fully reproducible.

Say which of those your scheduler offers, in its own documentation. Every built-in does, and Determinism is the contract those words refer to.

One trap that is not yours to fix

A pruned trial's objective value is its last intermediate report. If a study's reported metric gets worse as the resource grows, the trials your scheduler stopped early will beat the ones that ran to completion, and any sampler learning from that history is steered away from the good region.

That is worth knowing while writing a scheduler so that you do not try to correct it in the scheduler — the fix belongs in the objective, and Prune and schedule states it in full, with the measurement that made the point.

Where to go next

If you want to Go to
See the seam from the user's side Prune and schedule
See fork and revive in a real study Population-based training
Understand how the two seams relate Samplers and schedulers
Choose a trial's parameters instead Write a sampler
Ship it as a crate someone can depend on Publish a plugin crate
Read the trait, method by method Rust API