Search spaces¶
A search space is the set of configurations a study is allowed to consider: one entry per parameter, each saying what values that parameter may take. It is the most consequential thing you write. A sampler can only be as good as the space it searches, and the two most common ways to get a bad result out of an optimiser are a bound that excludes the answer and a scale that hides it.
This page is language-neutral: it names each concept and then how each language spells it.
The four kinds of parameter¶
There are exactly four, and they do not collapse into each other. A boolean is not an integer that happens to be 0 or 1, and a category is not a small integer — keeping them distinct is what lets a sampler treat each one correctly.
| Kind | What it holds | Rust | Python |
|---|---|---|---|
| Float | A real number in [low, high] |
suggest_f64(name, low..=high, scale) |
suggest_float(name, low, high, log=…, step=…) |
| Integer | A whole number in [low, high] |
suggest_int(name, low..=high, scale) |
suggest_int(name, low, high, log=…, step=…) |
| Categorical | One of a fixed, ordered list of labels | suggest_cat(name, &choices) |
suggest_categorical(name, choices) |
| Boolean | true or false |
suggest_bool(name) |
suggest_bool(name) |
Here are three of them in one objective, taken from a real example — a float width, an integer depth and a float learning rate:
/// Evaluates one configuration, returning `[error, cost]`.
///
/// Capacity buys accuracy and costs money; the learning rate buys accuracy for
/// free. A configuration that wastes its capacity on a bad learning rate is
/// dominated by one that does not, which is what gives the front a shape.
fn evaluate(ctx: &mut TrialCtx<'_>) -> Result<Outcome> {
let width = ctx.suggest_f64("width", 1.0..=64.0, Scale::Linear)?;
let depth = ctx.suggest_int("depth", 1..=8, Scale::Linear)?;
let lr = ctx.suggest_f64("lr", 0.001..=1.0, Scale::Linear)?;
// `f64::from` on an `i32` is lossless; the Python arm's `float(depth)` is the
// same widening, so both arms multiply the same two numbers.
let capacity = width * f64::from(i32::try_from(depth).unwrap_or(i32::MAX));
let off = lr - LR_STAR;
let error = 1.0 / (1.0 + capacity) + off * off;
Ok(vec![error, capacity].into())
}
def objective(trial: atune.Trial) -> list[float]:
"""Evaluates one configuration, returning ``[error, cost]``.
Capacity buys accuracy and costs money; the learning rate buys accuracy for
free. A configuration that wastes its capacity on a bad learning rate is
dominated by one that does not, which is what gives the front a shape.
"""
width = trial.suggest_float("width", 1.0, 64.0)
depth = trial.suggest_int("depth", 1, 8)
lr = trial.suggest_float("lr", 0.001, 1.0)
capacity = width * float(depth)
off = lr - LR_STAR
error = 1.0 / (1.0 + capacity) + off * off
return [error, capacity]
A categorical parameter stores the index of the chosen label, not the label itself, and the labels live once in the space — so a trial's record does not carry a copy of the vocabulary. What it does carry is the distribution it was drawn from, and the recorded choice list is compared element by element against the one you offer next time. Editing that list at all — renaming a label as much as adding or removing one — is therefore a different parameter as far as an existing study is concerned, and it is refused rather than reinterpreted.
Scale¶
Every numeric parameter is drawn on one of two scales, and picking the wrong one is the single most common way to waste a budget.
- Linear — uniform in the value. Correct for a width, a batch size, a number of layers.
- Logarithmic — uniform in the logarithm of the value. Correct for anything
whose ratio matters rather than whose difference does: a learning rate, a
regularisation coefficient, a discount's complement. On a linear scale, half
of a range from
1e-5to1e-1sits above0.05, and a learning-rate search spends half its budget there for no reason.
Rust names the scale explicitly (Scale::Linear, Scale::Log); Python takes
log=True. Two rules follow from what a logarithm is:
- A logarithmic parameter needs a strictly positive lower bound. Zero is not in the domain, and asking for it is rejected rather than clamped.
- A parameter cannot be both logarithmic and stepped. A grid in the value and a uniform draw in its logarithm describe different distributions, so combining them is refused.
Steps, and where the top of a stepped range goes¶
A step turns a continuous range into a grid, anchored at low. There is one
behaviour worth knowing before it surprises you: the upper bound is snapped
down onto the grid. A float parameter over 0.0..=11.0 with step 3.0 holds
the points 0, 3, 6, 9 — and records its high as 9, not 11, because 11 is not a
point the parameter can take. Anything else would leave the space claiming a
value it can never produce.
Two ways to write a space, and a third for configuration files¶
Define-by-run is the default and is what every example on this site uses.
There is no separate declaration: the first suggest_* call for a name is the
declaration. It creates that parameter's distribution, samples it, and records
both on the trial; a second call for the same name inside the same trial replays
the stored value rather than drawing again.
The consequence people like is that a space can depend on the trial. The consequence people forget is that a space defined this way is only known retrospectively — the study infers what the parameters were from what has been suggested so far.
A declared space is the other way, and it is Rust-only today. A struct with
one field per parameter and a #[derive(Space)] on it produces the space, the
decoding from an assignment and the encoding back, with the bounds written as
attributes next to the
fields they belong to. When a study is given a declared space up front, the
sampler works from the exact space rather than from an inference over history.
A string form exists for spaces that come from configuration rather than from
code — a TOML file, a command-line flag. One line of text per parameter, which is
what the CLI's --param flag and
Tune any program take; its
grammar is generated from the parser
that implements it, so what is written down and what is accepted cannot drift
apart.
All three routes lower to the same distributions, and none of them is a second implementation — which is what makes them interchangeable rather than merely similar. A space is not a second-class citizen for having arrived as text, and the choice between the three is about where the space is written: in the objective that consumes it, in a type you already have, or in a file someone else edits.
Conditional parameters¶
atune has no enabled_if. A parameter is not declared active only when another
parameter takes a particular value, and neither the derive attributes nor the
string form has a spelling for it.
What works instead is define-by-run: put the branch in the objective, and simply
do not call suggest_* for the arm you did not take. A trial that never
suggested momentum has no momentum, and the study accepts that a trial's
parameters are a subset of what other trials suggested. This is less expressive
than a declared conditional space — a sampler cannot reason about a branch it
only learns about after the fact — and it is the honest state of things today.
What a trial's parameters actually are¶
The parameters of one trial are a map from name to typed value, and the map is always iterated in sorted name order. That is not a formatting preference: several things downstream — the printed parity block, the derived per-parameter seeds, the on-disk record — would otherwise depend on the order a hash map happened to produce, which is exactly the kind of dependency that makes a "reproducible" run irreproducible. See Determinism.
Each value carries its kind with it, so a boolean true is never confused with
the integer 1, and a categorical index is never read as a number to do
arithmetic on.
When a space is wrong¶
Every route into a space validates it, and the validation happens in the
distribution constructors rather than in any one route — which is why a space
written as text, as attributes or by suggest_* is held to exactly the same
standard. What is worth carrying in your head is that a space can be wrong in
three distinguishable ways, and which error you get says where to look:
- The space itself is malformed — a bound that cannot be satisfied, a name
used twice, a scale that contradicts a bound. That is
Error::InvalidSpace, and its message names the parameter and what is wrong with it. - A value does not fit the space it is checked against:
Error::OutOfRange, naming the parameter. The space is fine; something handed it a value it cannot hold. - A parameter contradicts what the study already recorded under that name:
Error::Incompatible. Fewer edits trigger it than you might expect: a numeric parameter's bounds and its step may drift freely between trials, because a define-by-run objective is allowed to widen or narrow a range as it goes, and a value recorded under an earlier range stays valid even once it falls outside the current one. What is actually frozen is the parameter's kind (float, integer, categorical, boolean), its log flag, and — for a categorical — its exact choice list; change any of those under the same name and this is the error you get. The reference says exactly which edits survive and which do not.
A fourth case never becomes an error at all: a malformed #[space(...)]
attribute is a compile error
with a span, so the derive route rules out at build time what the other two can
only reject at run time.
Reading a boundary finding¶
A space can also be wrong in a fifth way that is never an error: every bound
is satisfiable and every trial is valid, but the range you picked was too
tight, and the sampler has been quietly telling you so one trial at a time.
atune doctor reads a study's completed
trials and reports this as a boundary finding, its most common one: the
good trials keep landing at one edge of a declared range, which usually means
the edge was guessed too tight rather than that the true optimum sits exactly
on it.
The finding names the crowded side — rendered lower or upper — and a
pressure: the fraction of the top trials sitting in the outer band on that
side, measured on the parameter's own scale (logarithmic for a log
parameter, linear otherwise). A pressure at or above one half is what gets
reported, and only when the best trial found so far itself sits in the band —
crowding without the leader is noise, not evidence; 1.0 means every one of
the top trials is in the band. For a plain range the fix is yours to make — atune doctor names the
parameter and the direction, and widening the range is an edit you make and
re-run, exactly like any other change to a search space, which numeric bound
drift already supports (see above). A range you
declared open closes that loop itself. Read a
doctor report walks through a full example
and the other findings alongside it.
Open search spaces¶
The boundary finding above has a standing answer: declare, up front, that a
bound is a guess. An open range starts at its seed exactly as a plain
range would, and may grow while the study runs — whenever the best trials
crowd one of its edges, that bound steps outward (doubling the span on a
linear scale, the ratio on a logarithmic one), until a brake says stop: an
explicit limit, an expansion cap, a run of failures in the newly added
region (which rolls the growth back), the zero line for a range that must
keep its sign, or the edge of what the type can represent.
The declaration rides any of the existing spellings — uniform(0, 10,
open=up, limit=..=200) in the DSL, the same keywords in a #[space(...)]
attribute, the Open builder in Rust, suggest_open_float in Python — and
around(c, times=n) declares "centered on c, no idea how far" without
picking hard bounds at all. The DSL
reference has the full grammar.
Three properties are worth carrying in your head, because they are what make an open space safe to hand to an overnight run:
- The seed is stated once. A policy is fixed when it is first declared; a resume or a second worker passing a different seed, side or limit is refused, exactly as any other contradiction with the stored study.
- Every trial keeps the range it was actually drawn from. Growth never rewrites history: a trial's record names its own bounds, so an analysis months later can tell which trials saw the wider world.
- Every decision is observable. A bound move or a brake is a study
event;
atune runprints one line per decision as it happens, andatune doctorexplains after the fact why a side stopped.
Two sampler families cannot follow a moving bound — one enumerates a
snapshot (Grid), the others keep their model in coordinates relative to a
fixed box (Cmaes, Dehb, Nsga2) — so a study that declares an open
space refuses them at create and the automatic choice simply never picks
them. A multi-objective study is refused with any open policy for the
same reason one step removed: Nsga2 is the only front-ranking built-in,
and it is one of the fixed-box samplers. Random, Qmc, Tpe, GpEi and
Carbs all recompute from history and work unchanged. One honest caveat: a bound only grows when the best
trials crowd it, and a pure Random sampler rarely concentrates its best
trials anywhere — under Random an open declaration is mostly dormant, and
a learning sampler such as Tpe is what makes it earn its keep. Under
several workers an open study is replayable rather than pre-determined; the
determinism contract states
exactly what that means.
Where to go next¶
| If you want to | Go to |
|---|---|
| Write your first one | Your first study |
| See who draws from the space | Samplers and schedulers |
| Give a space to a program from the shell | Tune any program |
| Read the string form's grammar | Search-space DSL |
| See an open range find a moved optimum | the paired open_range example (examples/rust, examples/python) |
| Read the derive's attribute grammar | #[derive(Space)] attributes |
| Look up an error you hit | Error reference |
| Find out which of your ranges are too tight | Read a doctor report |