The reinforcement-learning evidence¶
atune has features that a general-purpose tuner does not, and each of them is there because a published result says the ordinary approach fails on reinforcement-learning workloads. This page is that reasoning, with the citations, and with the boundary of what has actually been measured drawn explicitly.
It is an argument page, not a measurement page. Nothing here is a claim about atune's quality against another tool; for measurement see Benchmarks, and for capability-by-capability citations see Prior art.
Why RL is a different problem¶
An optimiser assumes that evaluating a configuration tells it how good the configuration is. In RL that assumption fails in four ways at once:
- Trials are noisy. The same configuration, the same code and the same budget give different answers under different seeds.
- Trials are expensive. A single evaluation is a training run, so budgets are tens of trials rather than thousands.
- The best setting moves during training. A learning rate or entropy coefficient that is right at the start is not right at the end.
- Failure is normal. Divergence and NaN are ordinary outcomes, not bugs in the search.
Every RL-oriented feature in atune is a response to one of those four.
What the literature establishes¶
Tuning works, and it is cheap relative to hand-tuning. Eimer, Lindauer and Raileanu (Hyperparameters in Reinforcement Learning and How To Tune Them, ICML 2023) report that a 64-run DEHB search beat an 810-run grid search — and that multi-fidelity and population-based methods, not plain Bayesian optimisation, are what survive the noise. atune's response is that DEHB and the population methods are built in rather than being plugins.
Configurations overfit the tuning seeds. This is the finding of the same paper that most tools ignore: a configuration selected on a set of seeds can be substantially worse on held-out seeds, so a search that reports its best tuning score is reporting the wrong number. atune's response is the multi-seed protocol — several tune seeds per trial, a disjoint test-seed set used only by a final re-evaluation stage, and a reported gap between the two. The RL tutorial shows that gap being produced on a runnable example.
Aggregation over seeds needs a robust statistic. Agarwal et al. (2021, the rliable / statistical-precipice work) argue that point estimates over a handful of runs mislead, and recommend the interquartile mean. atune offers the IQM as a built-in aggregate alongside the mean, the median and a spread-penalised mean — with one caveat the tutorial spells out: the IQM trims nothing at three seeds or fewer, so at small k it is exactly the mean.
Schedules beat constants. Population-based training (Jaderberg et al., 2017) and PB2 (Parker-Holder et al., NeurIPS 2020) work by forking a running trial and perturbing its parameters, which effectively learns a schedule that no static-configuration search can express. That is why atune's scheduler seam has pause, resume and fork semantics rather than only "prune", and why population-based training is a first-class path. The PBT-Zoo comparison (TMLR) is also where the critique of greedy exploitation comes from, which is why a variance-aware exploit option exists at all.
Early stopping works when the curve is informative, and misfires when it is not. The successive-halving line (Li et al., 2020, on ASHA and massively parallel search; Falkner et al., 2018, on BOHB) is what makes multi-fidelity search cheap. On a noisy RL curve, though, an early prune throws away good configurations, which is why atune's pruners come with noise-aware forms — a patience decorator, a median rule that fires only when the gap exceeds the spread, and a Wilcoxon test over per-seed samples. See Prune and schedule.
The landscape is noisy but low-dimensional. ARLBench (arXiv:2409.18827) reports that few hyperparameters carry most of the variance, which is what makes importance analysis worth computing on an RL study rather than a diagnostic curiosity.
In-run adaptation is a category, not a variant. ULTHO
(arXiv:2503.06101) adjusts hyperparameters
within a single training run using a clustered bandit, at a fraction of the
compute of a population method. A tuner can only do that if a decision costs
close to nothing, which is the requirement atune::OnlineTuner is built to.
Freeze-thaw needs pausing to exist first. The DyHPO (2022) and ifBO (ICML 2024) line of work resumes a paused trial when its extrapolated curve becomes attractive again. The prerequisite is a scheduler that can pause with a checkpoint and resume with typed state — the same machinery population-based training needs.
One number, and what it is not¶
Illustrative, and not reproducible from this repository
A cheap soft actor-critic recipe on CartPole, at a fixed budget, returned mean scores of roughly 183, 86 and 119 on seeds 0, 1 and 2 — same code, same configuration, same budget.
That measurement comes from an internal harness driving an external RL implementation, so it cannot be reproduced from this repository. It is here to show the size of the seed effect, not to say anything about atune's quality and not to compare atune with anything. Every number that does support a comparison lives on Benchmarks with the command that reproduces it.
Two configurations compared on one seed apiece, at that budget, are being compared on noise. That is the whole argument for the multi-seed protocol in one line.
What has not been measured here¶
Being explicit about this matters more on this page than anywhere else, because the literature above is other people's evidence and it is easy to let it stand in for our own:
- No RL benchmark has been run in this repository. atune's published measurements are on synthetic single-objective problems (Benchmarks). The RL-oriented schedulers are guarded by in-repository oracles and by an end-to-end study against an external agent implementation, and that study's numbers are not reproducible from this repository — so they are not published.
- The literature's results are about methods, not about this implementation. "DEHB beats grid search in an ICML study" is evidence that the method is worth having. It is not evidence about atune's DEHB.
- The seed-overfitting effect is reported, not re-measured. What this repository demonstrates is that the protocol produces the tune-versus-test gap on a runnable objective, which is a statement about the machinery.
Where to go next¶
| If you want to | Go to |
|---|---|
| The protocol as a lesson, on a runnable example | Tuning an RL agent |
| The protocol as a recipe | Multi-seed RL protocol |
| Schedules rather than constants | Population-based training |
| Pruning that survives noise | Prune and schedule |
| What ships, and behind which feature | Catalog |
| Measurements with reproduction commands | Benchmarks |