Skip to content

Use a multi-seed evaluation protocol

Goal: stop a lucky seed from winning your search.

Reinforcement learning — and anything else with a noisy objective — has a failure mode that looks like success. Score each configuration on one seed, take the best, and what you have found is partly the configuration and partly a favourable random draw. Re-run the winner on a fresh seed and some of your improvement evaporates.

The fix is standard practice: evaluate each configuration on several seeds, aggregate robustly, and hold out seeds you never tuned on. atune makes it a property of the study rather than something you build around it.

Declare the protocol

// `replicate_seed` is the seed *this replicate* must run under — common
// across trials inside a fan, disjoint from the test seeds. A reproducible
// objective takes all of its randomness from it and none from ambient
// entropy.
let objective = |ctx: &mut TrialCtx<'_>| -> Result<Outcome> {
    let x = ctx.suggest_f64("x", -5.0..=5.0, Scale::Linear)?;
    let distance = x - TARGET - seed_shift(ctx.replicate_seed());
    Ok((distance * distance).into())
};

let study = Study::builder()
    .parallelism(THREADS)
    .budget(Budget::trials(TRIALS))
    .seed_protocol(
        SeedProtocol::new()
            .with_tune_seeds(TUNE_SEEDS)
            .with_test_seeds(TEST_SEEDS)
            .with_aggregate(Aggregate::Iqm)
            // The default, written out because it is one of the protocol's
            // four decisions: trials share tuning seeds (common random
            // numbers), so comparisons subtract the seed's contribution.
            .with_pairing(Pairing::Paired),
    )
    .create(StudyConfig::new(STUDY).with_seed(SEED))?;

study.optimize(objective)?;

// Re-rank the top tune-ranked trials on the disjoint test seeds.
let report = study.reevaluate(objective, TOP_K)?;
study = atune.create_study(
    direction="minimize",
    n_seeds=TUNE_SEEDS,
    n_test_seeds=TEST_SEEDS,
    aggregate=AGGREGATE,
    seed=SEED,
    name=STUDY,
)
study.optimize(objective, n_trials=TRIALS, n_jobs=JOBS)

# Re-rank the top tune-ranked trials on the disjoint test seeds.
report = study.reevaluate(objective, top_k=TOP_K)

Four decisions, all on the study rather than in your objective. Rust spells all four on the SeedProtocol builder, as the listing shows; Python takes them as create_study keywords with lowercase string values:

  • n_seeds — how many tuning replicates each trial runs. Each replicate gets its own seed, derived from the study seed so the whole thing stays reproducible.
  • n_test_seeds — validation seeds, disjoint from the tuning ones. The objective does not tune directly on them, but if you choose or retune a candidate after inspecting these results, they have become selection data.
  • the aggregator (aggregate=) — how replicates collapse to one number: "iqm" (interquartile mean, Iqm in Rust) trims the outer quarter on each side when enough replicates are available; with three seeds, it equals the mean under atune's current trimming rule; "mean_minus_std" (MeanMinusStd) computes mean - risk_lambda * std, using the population standard deviation. A positive risk_lambda penalizes spread when maximizing; use a negative value when minimizing. "mean" and "median" are also available. (The CLI flag spells the risk-adjusted aggregate mean-minus-std.)
  • the pairing (pairing=) — whether trials share tuning seeds ("paired"/Paired, the default) or draw independent ones. Paired is a common-random-numbers design: it controls seed effects shared across configurations and can reduce comparison variance. Configuration-by-seed interactions remain, so the variance reduction depends on the objective.

Your objective does not change shape. It is called once per replicate and returns one number; the fan and the aggregation happen around it.

Re-evaluate on validation seeds

Tuning-set performance is optimistic by construction — you selected on it. Re-evaluation on disjoint validation seeds gives a useful check on that selection. The listing above ends with the call: study.reevaluate(objective, k).

reevaluate takes the top k trials by tuning score, runs them on the validation seeds, and reports both aggregates and their gap. Treat the gap as a diagnostic, not a guarantee that the ranking generalises. The test replicates are stored on the trials while their tuning objectives remain unchanged.

The validation ranking can differ from the tuning ranking. If you use it to choose a winner, report that choice as validation-selected; do not present the same validation results as an untouched final test. For an unbiased final performance claim after candidate selection or retuning, evaluate the chosen configuration on a separate seed set that has not informed any decision.

Keep the paired per-seed outcomes as well as aggregates. Paired seeds make configuration comparisons on shared seeds possible; showing those raw outcomes and their variation makes it easier to see whether an aggregate hides an unstable result.

Before spending a large budget, check that the objective uses the intended units and direction, and sanity-check a few outcomes against the environment or learner's expected scale. Track quality and compute cost separately. If you combine them into a penalized objective, document the penalty: it changes which configuration is optimal.

Keep experiment provenance

For each tuning run, each decision that selects or retunes a candidate, and the final evaluation, keep a consumer-owned record of the source revision and dirty patch identity; the resolved configuration and objective definition, including the metric window and transformations, units, direction, and any penalty; the named seeds and stage; backend, runtime, dependency versions, and relevant hardware; checkpoint artifact identity or digest and what resume restores; and raw per-seed objective and compute outcomes with their units. Atune provides generic user-owned study and trial attributes, but does not automatically capture this provenance. See Study and trial for those attribute locations.

Make the example actually show something

Two ways to build a multi-seed setup that runs correctly and demonstrates nothing. Both were hit while writing this project's own example.

The seed must interact with the configuration. Under paired seeding the tuning seeds are common across trials, so noise that depends on the seed alone adds the same offset to every trial's aggregate — every overfit gap identical, and the held-out ranking a carbon copy of the tuning ranking. Measured: all ten re-evaluation entries reported the same gap to the last digit. Real noise interacts with the configuration — a learning rate that is good on one initialisation and bad on another — so a demonstration has to as well.

Derive the noise from the seed atune hands you, not from a global generator. A process-global random number generator that nothing seeds is not reproducible even in one language, and two languages' generators never agree. The example mixes the replicate seed arithmetically, which is why its Rust and Python arms agree bit-for-bit.

What it costs

Multiplicatively more compute: n_seeds replicates per trial, plus n_test_seeds × k for the re-evaluation. Three tuning seeds turn a 100-trial study into 300 evaluations.

Choose the seed count against both noise and compute cost. Three seeds improve replication over a single seed, but atune's IQM is the arithmetic mean at that count, so it does not trim outliers. More seeds increase evaluation cost and allow the IQM's trimming rule to take effect; fewer trials may then fit the same budget. Report the count and per-seed variation so readers can interpret the aggregate.

Next