Skip to content

Why atune

Hyperparameter optimisation is a solved problem when a trial is a model.fit() call: the search costs seconds, the trial costs hours, and the tuner's own overhead is noise. atune exists for the cases where that stops being true — where the training loop is native code, where a trial takes a millisecond, where the tuner has to live inside the loop, or where the whole thing has to ship as one binary.

This page states what atune does that is unusual, as capability rather than comparison. It makes no claim that atune is better than anything; where a comparison exists it is a measurement or a citation, and it lives on Benchmarks and Prior art instead.

Nothing is released yet

Every capability below is in the tree and gated by a test. None of it is on crates.io or PyPI — see Install for how to build from a clone.

The unusual things

The core is compiled, and it runs where Python cannot

The model layer — studies, trials, spaces, the four plugin seams, the stateless samplers — is a Rust crate with a small dependency set that compiles to wasm32-unknown-unknown. That is not an aspiration: the release gate runs cargo check -p atune_core --target wasm32-unknown-unknown --no-default-features on every change, so anything that reaches for a clock, a file or a thread inside the core fails the build. Everything that needs an operating system lives behind a cargo feature.

The consequence a user notices is that the objective is a function in their own process. No subprocess, no interpreter in the hot path, and trials that run on real threads rather than under a global lock.

Determinism is a contract, not a best effort

With a stateless sampler, the parameters of trial n are a pure function of the study seed, n and the search space. Changing the thread count changes the wall clock and nothing else. Determinism states the contract exactly, including where it stops holding.

What makes it more than a promise is how it is checked. The library's own tests assert the same parameter sequence across 1, 4 and 8 threads; on top of that, a cross-language parity gate runs the Rust and Python versions of seven examples and requires their results to be equal to the bit — same trial number, same parameters, same objective value compared as raw f64 bits. A sampler change that alters the answer in one language and not the other cannot land.

Search spaces come from a type, or from a file

Three ways to declare a space, one engine underneath:

  • #[derive(Space)] on a plain configuration struct — ranges, log scales, steps and choices as attributes, checked at compile time.
  • A TOML overlay — a [space] table added to a program's existing configuration file, so a TOML-configured program becomes tunable with no code change at all. The grammar is on the space DSL reference.
  • Define-by-run — trial.suggest_f64("lr", …) inside the objective, for spaces whose shape depends on earlier draws.

Parameters are keyed by name and carry a typed value model rather than being collapsed into one numeric union, which is what lets a study be resumed and correlated across processes.

A wrong range can fix itself while the study runs

Every declaration above accepts one more word: open. A range declared open starts at its seed exactly as a closed one would, and grows during the study whenever the best trials crowd one of its edges — bounded by an explicit limit, an expansion cap, a failure veto that rolls a bad expansion back, and the zero line for a range that must keep its sign. The wrong guess that normally costs an overnight run gets corrected at the trial where the evidence appears, not the morning after.

What makes this shippable rather than a heuristic: every trial's record keeps the range it was actually drawn from, every decision is a study event (atune run prints one line per move, atune top logs each bound movement, and the GUI charts the bounds over time), a single worker with a fixed seed replays the growth decisions identically, and the configurations that cannot follow a moving bound are refused at create rather than silently degraded — Grid and the population/covariance samplers, and a multi-objective study under any open policy. For a study you would rather only diagnose, atune doctor reports the same evidence as advice instead. The concept page has the full contract.

The reinforcement-learning protocol is part of the study

RL breaks the assumption that evaluating a configuration tells you how good it is: the same configuration under a different seed is a different answer. atune makes the remedy a property of the study rather than something built around it — several tune seeds per trial with paired common random numbers, per-seed results kept rather than averaged away, a choice of aggregate including the interquartile mean, and a disjoint set of test seeds used by nothing until a final re-evaluation stage. The tune-versus-test gap it produces is the output that tells you whether the search overfit its seeds. Multi-seed RL protocol is the recipe, and RL evidence is why.

Schedulers pause, fork and mutate — they do not only prune

A pruner can stop a trial. A scheduler can also suspend one, resume it later, and fork a copy with perturbed parameters, which is what population-based training needs and what turns "tune a constant" into "learn a schedule". Both kinds live behind one seam; the catalogue of what ships is on the reference, and Prune and schedule covers the choice.

A tuner small enough to live inside the training loop

atune::OnlineTuner adapts hyperparameters between updates of a single run instead of across trials. It is a different shape of tuning from a study, and it is in the same library; whether a particular loop meets its latency budget depends on the objective and deployment, not on an unmeasured promise here.

One core, three surfaces, and a way out

The Rust library, the Python package and the atune binary are three front ends over the same core, and none of them is the "real" one. The CLI tunes a program that reads environment variables and prints a number, in any language, with no bindings at all. Studies can be exported to and imported from Optuna's journal format, so an existing Optuna dashboard keeps working and adopting atune is not a one-way door.

What atune is not

Stated as plainly as the capabilities, because a tool that claims no boundary usually has one anyway:

  • Not a numerical-optimisation library. Gradient-based and local optimisation belong elsewhere; atune implements the algorithms hyperparameter search needs and no more.
  • Not a cluster manager. Distribution is mediated by shared storage and worker processes. There is no scheduler daemon, and launching onto a cluster is somebody else's adapter.
  • Not a neural-architecture-search framework. Categorical parameters and define-by-run branching cover simple architecture choices; supernets and weight sharing are out of scope.
  • Not an experiment tracker. There is a study viewer, not a metrics platform; exporters are the integration point.
  • Not an AutoML pipeline system. No feature engineering, no model-selection graphs.

Where to go next

If you want to Go to
The evidence behind the RL choices RL evidence
Which tool has which capability, with citations Prior art
Measurements, with the command that reproduces each Benchmarks
To start using it Install
To understand the vocabulary first Studies and trials
To plug in your own sampler Extending