Why atune¶
Hyperparameter optimisation is a solved problem when a trial is a
model.fit() call: the search costs seconds, the trial costs hours, and the
tuner's own overhead is noise. atune exists for the cases where that stops being
true — where the training loop is native code, where a trial takes a
millisecond, where the tuner has to live inside the loop, or where the whole
thing has to ship as one binary.
This page states what atune does that is unusual, as capability rather than comparison. It makes no claim that atune is better than anything; where a comparison exists it is a measurement or a citation, and it lives on Benchmarks and Prior art instead.
Nothing is released yet
Every capability below is in the tree and gated by a test. None of it is on crates.io or PyPI — see Install for how to build from a clone.
The unusual things¶
The core is compiled, and it runs where Python cannot¶
The model layer — studies, trials, spaces, the four plugin seams, the stateless
samplers — is a Rust crate with a small dependency set that compiles to
wasm32-unknown-unknown. That is not an aspiration: the release gate runs
cargo check -p atune_core --target wasm32-unknown-unknown --no-default-features
on every change, so anything that reaches for a clock, a file or a thread inside
the core fails the build. Everything that needs an operating system lives behind
a cargo feature.
The consequence a user notices is that the objective is a function in their own process. No subprocess, no interpreter in the hot path, and trials that run on real threads rather than under a global lock.
Determinism is a contract, not a best effort¶
With a stateless sampler, the parameters of trial n are a pure function of the study seed, n and the search space. Changing the thread count changes the wall clock and nothing else. Determinism states the contract exactly, including where it stops holding.
What makes it more than a promise is how it is checked. The library's own tests
assert the same parameter sequence across 1, 4 and 8 threads; on top of that, a
cross-language parity gate runs the Rust and Python versions of seven
examples and requires their results to be equal to the bit — same trial
number, same parameters, same objective value compared as raw f64 bits. A
sampler change that alters the answer in one language and not the other cannot
land.
Search spaces come from a type, or from a file¶
Three ways to declare a space, one engine underneath:
#[derive(Space)]on a plain configuration struct — ranges, log scales, steps and choices as attributes, checked at compile time.- A TOML overlay — a
[space]table added to a program's existing configuration file, so a TOML-configured program becomes tunable with no code change at all. The grammar is on the space DSL reference. - Define-by-run —
trial.suggest_f64("lr", …)inside the objective, for spaces whose shape depends on earlier draws.
Parameters are keyed by name and carry a typed value model rather than being collapsed into one numeric union, which is what lets a study be resumed and correlated across processes.
A wrong range can fix itself while the study runs¶
Every declaration above accepts one more word: open. A range declared open
starts at its seed exactly as a closed one would, and grows during the
study whenever the best trials crowd one of its edges — bounded by an
explicit limit, an expansion cap, a failure veto that rolls a bad expansion
back, and the zero line for a range that must keep its sign. The wrong guess
that normally costs an overnight run gets corrected at the trial where the
evidence appears, not the morning after.
What makes this shippable rather than a heuristic: every trial's record keeps
the range it was actually drawn from, every decision is a study event
(atune run prints one line per move, atune top logs each bound movement,
and the GUI charts the bounds over time), a single worker with a fixed seed
replays the growth decisions identically, and the configurations that cannot
follow a moving bound are refused at create rather than silently degraded —
Grid and the population/covariance samplers, and a multi-objective study
under any open policy. For a study you would rather only diagnose,
atune doctor reports the same evidence
as advice instead. The concept page has the
full contract.
The reinforcement-learning protocol is part of the study¶
RL breaks the assumption that evaluating a configuration tells you how good it is: the same configuration under a different seed is a different answer. atune makes the remedy a property of the study rather than something built around it — several tune seeds per trial with paired common random numbers, per-seed results kept rather than averaged away, a choice of aggregate including the interquartile mean, and a disjoint set of test seeds used by nothing until a final re-evaluation stage. The tune-versus-test gap it produces is the output that tells you whether the search overfit its seeds. Multi-seed RL protocol is the recipe, and RL evidence is why.
Schedulers pause, fork and mutate — they do not only prune¶
A pruner can stop a trial. A scheduler can also suspend one, resume it later, and fork a copy with perturbed parameters, which is what population-based training needs and what turns "tune a constant" into "learn a schedule". Both kinds live behind one seam; the catalogue of what ships is on the reference, and Prune and schedule covers the choice.
A tuner small enough to live inside the training loop¶
atune::OnlineTuner adapts hyperparameters between updates of a single run instead of
across trials. It is a different shape of tuning from a study, and it is in the
same library; whether a particular loop meets its latency budget depends on the
objective and deployment, not on an unmeasured promise here.
One core, three surfaces, and a way out¶
The Rust library, the Python package and the atune binary are three front ends
over the same core, and none of them is the "real" one. The CLI tunes a program
that reads environment variables and prints a number, in any language, with no
bindings at all. Studies can be exported to and imported from Optuna's
journal format, so an existing Optuna dashboard keeps working and adopting
atune is not a one-way door.
What atune is not¶
Stated as plainly as the capabilities, because a tool that claims no boundary usually has one anyway:
- Not a numerical-optimisation library. Gradient-based and local optimisation belong elsewhere; atune implements the algorithms hyperparameter search needs and no more.
- Not a cluster manager. Distribution is mediated by shared storage and worker processes. There is no scheduler daemon, and launching onto a cluster is somebody else's adapter.
- Not a neural-architecture-search framework. Categorical parameters and define-by-run branching cover simple architecture choices; supernets and weight sharing are out of scope.
- Not an experiment tracker. There is a study viewer, not a metrics platform; exporters are the integration point.
- Not an AutoML pipeline system. No feature engineering, no model-selection graphs.
Where to go next¶
| If you want to | Go to |
|---|---|
| The evidence behind the RL choices | RL evidence |
| Which tool has which capability, with citations | Prior art |
| Measurements, with the command that reproduces each | Benchmarks |
| To start using it | Install |
| To understand the vocabulary first | Studies and trials |
| To plug in your own sampler | Extending |