Tuning an RL agent¶
Reinforcement learning breaks the assumption every hyperparameter optimiser is built on: that evaluating a configuration tells you how good it is. Run the same agent, the same configuration and the same budget under three different seeds and you get three different answers, sometimes by a factor of two. A tuner that takes one of those numbers at face value will confidently hand you the configuration that got the luckiest seed.
This page is about the machinery atune has for that problem — a multi-seed evaluation protocol with a held-out test stage — and about reading its output honestly.
What you will actually run here
The program on this page is not an agent. It is a small deterministic objective whose optimum moves with the seed, which is the property of an RL environment that matters here, making it a compact stand-in for this tutorial.
This repository has no runnable RL example: the internal harness that drives a real agent is unpublished and needs a second project built alongside it. Rather than print a listing you cannot execute, this page teaches the protocol on something you can, and then says what changes when the objective is a real training run.
Why one seed is not enough¶
The concrete version of the problem, from this project's own RL work: a soft actor-critic recipe on CartPole, at a fixed budget, returned mean scores of roughly 183, 86 and 119 on seeds 0, 1 and 2. Same code, same configuration, same budget. Any two configurations compared on one seed apiece at that budget are being compared on noise.
That measurement comes from an internal harness driving an external RL implementation, so it is not reproducible from this repository alone — treat it as an illustration of the size of the effect rather than as a benchmark. The effect itself is not controversial; it is why the RL literature reports distributions over seeds rather than single runs.
The protocol¶
atune makes the fix a property of the study rather than something you build around it. A study can declare:
| Setting | What it does |
|---|---|
| Tune seeds, k | Each trial is evaluated k times, once per seed. The trial's objective — the one number the sampler ranks — is an aggregate over those k results. Per-seed results are kept, never lost to the aggregate. |
| Test seeds, m | A disjoint set, used by nothing during the search. They exist for the final stage only, so a configuration is never tested on a seed it was tuned on. |
| Pairing | paired (the default) uses the same tune seeds for every trial — common random numbers, which removes seed luck from the comparison between trials. independent gives each trial its own. |
| Aggregate | mean, median, iqm (the interquartile mean, robust to one lucky or unlucky seed) or a mean penalised by the spread. |
One detail about the interquartile mean that is easy to trip over: it trims
floor(k/4) results from each end, so for k of 3 or fewer it trims nothing and
is exactly the mean. Asking for iqm with three tune seeds is not wrong, but it
is not yet robust either.
Then, after the search, a re-evaluation stage takes the best few trials by their tune score and runs them again on the held-out seeds. Here is a study with all of that switched on:
// `replicate_seed` is the seed *this replicate* must run under — common
// across trials inside a fan, disjoint from the test seeds. A reproducible
// objective takes all of its randomness from it and none from ambient
// entropy.
let objective = |ctx: &mut TrialCtx<'_>| -> Result<Outcome> {
let x = ctx.suggest_f64("x", -5.0..=5.0, Scale::Linear)?;
let distance = x - TARGET - seed_shift(ctx.replicate_seed());
Ok((distance * distance).into())
};
let study = Study::builder()
.parallelism(THREADS)
.budget(Budget::trials(TRIALS))
.seed_protocol(
SeedProtocol::new()
.with_tune_seeds(TUNE_SEEDS)
.with_test_seeds(TEST_SEEDS)
.with_aggregate(Aggregate::Iqm)
// The default, written out because it is one of the protocol's
// four decisions: trials share tuning seeds (common random
// numbers), so comparisons subtract the seed's contribution.
.with_pairing(Pairing::Paired),
)
.create(StudyConfig::new(STUDY).with_seed(SEED))?;
study.optimize(objective)?;
// Re-rank the top tune-ranked trials on the disjoint test seeds.
let report = study.reevaluate(objective, TOP_K)?;
study = atune.create_study(
direction="minimize",
n_seeds=TUNE_SEEDS,
n_test_seeds=TEST_SEEDS,
aggregate=AGGREGATE,
seed=SEED,
name=STUDY,
)
study.optimize(objective, n_trials=TRIALS, n_jobs=JOBS)
# Re-rank the top tune-ranked trials on the disjoint test seeds.
report = study.reevaluate(objective, top_k=TOP_K)
The objective¶
/// Where this seed puts the bowl's minimum, relative to [`TARGET`].
///
/// A `SplitMix64` finalizer over `seed`, its top 53 bits divided by 2^53 to give a
/// fraction in `[0, 1)`, mapped onto `±SHIFT_SPREAD / 2`. Every step is exact —
/// wrapping 64-bit integer arithmetic, then one division by a power of two — so
/// the Python arm's unbounded-integer version produces the same bits rather than
/// approximately the same number. `z >> 11` leaves 53 bits, which is exactly what
/// an `f64` can hold without rounding.
fn seed_shift(seed: u64) -> f64 {
let mut z = seed.wrapping_add(0x9E37_79B9_7F4A_7C15);
z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
z ^= z >> 31;
#[allow(
clippy::cast_precision_loss,
reason = "`z >> 11` is 53 bits, which an f64 holds exactly — that is the point of the shift"
)]
let unit = (z >> 11) as f64 / TWO_POW_53;
SHIFT_SPREAD * (unit - 0.5)
}
def seed_shift(seed: int) -> float:
"""Where this seed puts the bowl's minimum, relative to ``TARGET``.
A SplitMix64 finalizer over ``seed``, its top 53 bits divided by 2^53 to give
a fraction in ``[0, 1)``, mapped onto ``±SHIFT_SPREAD / 2``. Every step is
exact — wrapping 64-bit integer arithmetic, then one division by a power of
two — so the Rust arm's ``u64`` version produces the same bits rather than
approximately the same number. The ``& MASK64`` masks are what make Python's
unbounded integers wrap the way ``u64`` does.
"""
z = (seed + 0x9E3779B97F4A7C15) & MASK64
z = ((z ^ (z >> 30)) * 0xBF58476D1CE4E5B9) & MASK64
z = ((z ^ (z >> 27)) * 0x94D049BB133111EB) & MASK64
z ^= z >> 31
unit = (z >> 11) / TWO_POW_53
return SHIFT_SPREAD * (unit - 0.5)
def objective(trial: atune.Trial) -> float:
"""Evaluates one replicate of one configuration.
``replicate_seed`` is the seed *this replicate* must run under — common
across trials inside a fan, disjoint from the test seeds. A reproducible
objective takes all of its randomness from it and none from ambient entropy.
"""
x = trial.suggest_float("x", -5.0, 5.0)
distance = x - TARGET - seed_shift(trial.replicate_seed)
return distance * distance
Two rules are baked into that function, and both are worth copying into your own:
- Derive the seed's effect arithmetically from the seed atune hands you. Not from a language random-number generator: a process-global generator that nothing seeds is not reproducible even in one language, and two languages' generators never agree. Here the seed goes through an integer mixer and one exact division, so both arms of the example compute identical bits.
- The seed must interact with the configuration. Under paired seeding the tune seeds are common across trials, so noise that depended on the seed alone would add the same offset to every trial's aggregate — every overfit gap identical, and the test ranking a carbon copy of the tune ranking. Here each seed moves the location of the optimum, which is what an RL seed actually does: every seed's environment has its own best configuration.
Run it, and read the gap table¶
cargo run -p atune --example multiseed, or python examples/python/multiseed.py.
Sixty trials, three tune seeds, two test seeds, aggregated with the interquartile mean, and the top ten re-evaluated. The interesting part of the output is the last column (abridged to six of the ten rows):
| Trial | Tune value | Test value | Gap |
|---|---|---|---|
| 0 | +0.514165 | +0.417570 | −0.096594 |
| 20 | +0.537388 | +0.403819 | −0.133568 |
| 56 | +0.538159 | +0.403503 | −0.134656 |
| 30 | +0.581906 | +0.835213 | +0.253307 |
| 36 | +0.583446 | +0.838390 | +0.254944 |
| 9 | +0.626486 | +0.923125 | +0.296639 |
Trial 0 won the search. Trial 56 — third by tune score, and within a thousandth of the runner-up — is the best configuration on seeds it had never seen. Trials 30, 36 and 9 all looked competitive during tuning and lost a quarter of a point or more when tested; that positive gap is overfitting to the tuning seeds, measured rather than suspected.
With one seed per trial none of this column exists, and trial 0 is simply "the answer".
Reading the answer¶
There is a trap here, and it is deliberate rather than accidental:
After re-evaluation, the study's best_trial is still the tune-ranked best.
The re-evaluation does not overwrite any trial's objective — doing so would
corrupt the history the sampler learned from. The test-ranked winner is on the
report the re-evaluation returns, and that is the study's honest answer. When
the two disagree — as they do above — the search overfit its tuning seeds, and
that disagreement is the signal the whole protocol exists to produce.
So: read best_trial if you want to know what the search converged on, and read
the report's best if you want to know what to ship.
What changes with a real agent¶
The protocol is the same. Three things around it are different:
- The objective is a subprocess, not a function. A training run is another program, and atune drives it by substituting parameters into its command line or environment and reading its output. Tune any program is that mechanism, and it does not care what language the agent is in.
- Fidelity is measured in environment steps. Declare the resource unit, and a successive-halving ladder's rungs mean what an RL practitioner expects them to mean rather than "however many times the objective happened to report".
- Pruning has to survive the noise. An early prune on a noisy curve throws away good configurations. The remedies in the catalogue are a patience decorator that suppresses pruning until a trial has genuinely stopped improving, and a median pruner that only prunes when the gap exceeds the spread of its competitors. Note also that pruning does not happen inside a multi-seed fan — reports are collected during the fan and the scheduler sees the trial when it ends.
Where to go next¶
| If you want to | Go to |
|---|---|
| The protocol as a recipe rather than a lesson | Multi-seed RL protocol |
| Stop weak trials early without being fooled by noise | Prune and schedule |
| Evolve hyperparameters during training | Population-based training |
| Tune an agent written in another language | Tune any program |
| The evidence behind atune's RL choices | RL evidence |
| Understand what the seeds guarantee | Determinism |