Skip to content

Samplers and schedulers

Two plugin seams decide how a study spends its budget, and they act at different moments:

  • A sampler answers what should the next trial try? It runs before a trial starts, and it sees the trials that have finished.
  • A scheduler answers what should happen to this trial, now? It runs while a trial is running, each time that trial reports an intermediate result.

Neither knows about the other. A study always has one of each; if you name neither, you get random search and a scheduler that never intervenes.

The sampler

A sampler is asked for the parameters of trial n and returns them. What it may consult in doing so divides the catalogue into three grades, and the grade is the thing to know about a sampler before you use it:

Grade Meaning Samplers
Stateless The draw is a pure function of the study seed, the trial number and the space. Nothing else. Random, Grid, Qmc
History-dependent, no stored state Recomputes a model from the finished trials on every ask; keeps nothing between calls. Tpe, GpEi, Carbs
Stateful Carries a population or a bracket across asks, persisted so a resumed study continues rather than restarts. Dehb, Nsga2, Cmaes

What each one is — its own one-line description, and the cargo feature it needs if it needs one — is the sampler catalogue, which is generated from the facade's re-export list, so a sampler that ships is a sampler that appears there. Three things it cannot tell you, because they are judgements rather than facts about a type:

  • Random is the default, and it is the baseline you have to beat. A model-based sampler that does not beat uniform search on your problem is not earning its keep, and running the comparison is the only way to find out.
  • Grid is exhaustive, so it ends. It hands out each point of a finite space once, and when there are none left the study stops with a space-exhausted error. That is a normal end rather than a failure: the search is over.
  • Qmc is stateless without being memoryless. The n-th Sobol point is a pure function of n, so it belongs in the first grade — but it is placed with the previous n − 1 in mind, and that is where its lower-discrepancy coverage comes from. Statelessness is a property of what a sampler reads, not of how much it knows.

One constraint cuts across the grades: a study that declares an open search space refuses Grid, Cmaes, Dehb and Nsga2 at create — Grid enumerates a snapshot of the space, the other three keep their model in coordinates relative to a fixed box, and neither can follow a bound that grows mid-study. The samplers that recompute from history each ask (Random, Qmc, Tpe, GpEi, Carbs) work unchanged.

There is also a helper that picks one for you from the space, the budget and the directions — multi-objective goes to Nsga2, a small finite space to Grid, a tiny budget to Qmc, a categorical-heavy space to Tpe, and so on. It is a pure function with no ambient state and no randomness, and it is opt-in: the default sampler stays Random whether or not you know the helper exists.

The scheduler

A scheduler is handed each intermediate report and answers with one of four decisions:

Decision Effect
Continue Nothing happens; the trial runs on.
Prune The trial stops now and is recorded as pruned.
Pause The trial suspends, holding a checkpoint reference, and may be resumed later.
Fork A new trial is created from this one's parameters and checkpoint. The reporting trial is not stopped.

Which of those four a scheduler is capable of answering is the fastest way to sort the catalogue, because it decides what the scheduler can do to your budget. The scheduler catalogue carries each one's description and, for the ones that need it, the cargo feature; the choice between them is this:

Family Answers with Reach for it when
Comparative stopping — MedianPruner, AshaPruner, HyperbandPruner Prune A trial can be judged against its peers at the same step. MedianPruner needs no ladder and starts working as soon as a few trials have passed the step in question; the two halving pruners want a rung ladder and pay for it with a prune rate you can predict from the reduction factor instead of from the shape of your metric. Hyperband is Asha hedged — several ladders at once, so a badly chosen first rung costs a fraction of the budget rather than all of it.
Statistical stopping — WilcoxonPruner Prune The reports are paired independent observations of the same quantity — several seeds, several folds, several evaluation episodes — and not successive points on one learning curve. Handing it a learning curve is the one way to misuse it.
Population — Pbt, Pb2 Fork Budget should be redirected rather than reclaimed. An underperformer's report copies a top performer's parameters and checkpoint, with the scalar ones perturbed, and nothing is pruned — so the population has to be large enough for "top" to mean something.
Curve extrapolation — FreezeThaw Pause A trial that looks bad early might still win. It fits each curve, suspends the ones whose asymptote is not competitive, and wakes the most promising later — reversible, where a prune is not.
Neither — NopScheduler, Patient — NopScheduler is the default and never intervenes. Patient is a decorator: it wraps another scheduler and swallows its prunes until a trial has gone n reports without improving on itself, which is the retrofit for a metric noisy enough to dip.

Why a pruning example's curve falls

This is the one rule to internalise before writing a pruned objective, and getting it wrong inverts the feature.

A pruned trial's objective value is its last intermediate value. It is not discarded, and it is not treated as infinitely bad — successive halving depends on that, and so does every sampler learning from a pruned trial.

Now suppose your reported metric rises with the resource — a cumulative reward, say, or an accuracy under a minimisation study. A trial killed at step 1 records a small number. A trial that survived to step 27 records a large one. Under minimisation the killed trial then beats every trial that ran to completion, the study's best trial is one that was stopped for being bad, and a model-based sampler that learns from those records is steered away from the good region rather than towards it.

That is not hypothetical. An earlier version of this repository's pruning example reported value = x * step; the best trial in the study was one pruned at step 1, and TPE converged on x ≈ 5 when the optimum was x = 0. The example argued against the feature it documented.

The fix is not a special case in the pruner. It is to report the quantity you are actually minimising, on a curve that improves with the resource — a loss that decays towards a floor, which is the shape a real learning curve has. Here is the corrected example's study, with the ladder and the sampler both set explicitly:

let study = Study::builder()
    .parallelism(THREADS)
    .budget(Budget::trials(TRIALS))
    // Successive halving over a 1..27 ladder: rungs at 1, 3 and 9.
    .scheduler(Arc::new(AshaPruner::new(
        u64::from(MIN_RESOURCE),
        u64::from(MAX_RESOURCE),
        REDUCTION_FACTOR,
    )?))
    .sampler(Arc::new(Tpe::new()))
    .create(StudyConfig::new(STUDY).with_seed(SEED))?;

study.optimize(|ctx| {
    let x = ctx.suggest_f64("x", LOW..=HIGH, Scale::Linear)?;
    for step in MIN_RESOURCE..=MAX_RESOURCE {
        // `report` returns `Error::TrialPruned` when the scheduler prunes;
        // letting `?` propagate it is the idiomatic path. The trial is then
        // recorded `Pruned` and keeps the intermediate it last reported —
        // which is why the curve has to *fall* (see the module header).
        ctx.report(u64::from(step), &[learning_curve(x, step)])?;
    }
    // A trial that survives the ladder is worth its converged loss.
    Ok(learning_curve(x, MAX_RESOURCE).into())
})?;
study = atune.create_study(
    direction="minimize",
    # Successive halving over a 1..27 ladder: rungs at 1, 3 and 9.
    scheduler=atune.schedulers.Asha(
        MIN_RESOURCE,
        MAX_RESOURCE,
        reduction_factor=REDUCTION_FACTOR,
    ),
    sampler=atune.samplers.Tpe(),
    seed=SEED,
    name=STUDY,
)
study.optimize(objective, n_trials=TRIALS, n_jobs=JOBS)

If your metric genuinely rises with the resource, minimise its negation, or maximise it and let the direction do the work. Do not leave the pruner to guess.

What a scheduler can and cannot see

Two limits are worth knowing, because they explain behaviour that otherwise looks like a bug.

The snapshot a scheduler is handed does not contain the trial that is reporting. It was taken when that trial was created. So a scheduler that needs the reporting trial's own history — a patience counter, a paired test, a curve fit — keeps that bookkeeping itself rather than reading it back.

During a multi-seed fan, reports are captured rather than streamed. A trial that evaluates its configuration under several seeds does not consult the scheduler in the middle of the fan, so there is no per-replicate pruning. Once every required replicate has reported the same step, atune deterministically aggregates that step and applies one scheduler decision to the whole logical trial. The finished trial's per-seed results are also visible when it ends. See Multi-seed RL protocol.

Determinism, honestly

The two seams have different standings, and every implementation states its own:

  • A stateless sampler is pre-determined. The parameters of trial n depend on the study seed, n and the space, and on nothing else — not on the worker count, not on completion order.
  • Everything else is replayable, not pre-determined. A history-dependent sampler and every real scheduler decide from the set of trials that have finished, and under parallelism that set depends on thread scheduling. Under a single worker with a fixed seed they are fully reproducible; under many workers the recorded history still explains the run exactly, but a rerun may differ.

This is why the examples on this site that use a pruner or a population run one worker. Determinism is the whole story.

Where to go next

If you want to Go to
Stop losing trials early, as a recipe Prune and schedule
Fork the winners mid-run Population-based training
Optimise two things at once Multi-objective and constraints
Write your own sampler or scheduler Extending: sampler, scheduler
Look one up by name, with its own description Catalog: samplers, schedulers
Find the feature flag one of these needs Feature reference