Determinism¶
Run a study twice with the same seed and you get the same study. That sentence is doing a lot of work, so this page states exactly what it promises, shows how it is asserted, and — at least as importantly — says where it stops.
The contract¶
For a stateless sampler, the parameters of trial number n are a pure function of the study seed, n, and the search space.
Not of the wall clock. Not of the thread interleaving. Not of the order trials finish in. Not of how many workers ran. A sixteen-thread study with a fixed seed and a fixed trial budget evaluates exactly the same set of configurations as a single-threaded one.
One declaration opts out of that sentence: an open search space. A
parameter declared open derives its live range from the study's own history,
so the space stops being an input you supplied and becomes a replayed
quantity — and under parallelism sixteen threads can widen at a different
trial than one thread does, then go on to evaluate a genuinely different set.
A single worker with a fixed seed remains fully reproducible; several workers
are replayable (the table below), and the policy's limit and expansion
cap are what bound how far two runs can diverge. This qualifies the space,
not any sampler — Random plus open is replayable even though Random
itself is stateless.
Two things the contract deliberately does not say. It does not say the trials finish in the same order — they do not, and that is fine. And it does not say the best value at any given instant is the same — only that the same configurations get evaluated, so the best value at the end is.
This is a test, not an aspiration. The library's determinism test builds the same 64-trial study at one, four and eight workers, collects the whole trial-number-to-parameters map from each, and asserts three things: that the numbers are contiguous and zero-based, that the parallel map equals the sequential one, and that repeating the four- and eight-worker runs reproduces it again — the worker count is not an input. A neighbouring test changes the seed by one and requires the result to differ, so the assertion cannot pass vacuously.
How a seed becomes a draw¶
There is one root of randomness — the study seed — and everything descends from it by a pure function:
- The study seed and the trial number are mixed, together with a stream tag, into a per-trial seed. The streams are separate: the sampler's randomness, the objective's, the scheduler's and each replicate of a multi-seed fan all get their own, so consuming randomness in one never shifts another.
- For each parameter, that per-trial seed is mixed again with the parameter's name.
- That final seed initialises a ChaCha8 generator, and the draw comes from it.
Three consequences fall out, and each of them is the reason for a design choice elsewhere on this site:
- The trial number is the key, never the trial id. Numbers are contiguous, per study, and assigned race-safely by storage; ids are opaque and may have gaps. Seeding on the id would make the answer depend on what else the storage happened to contain.
- Adding a parameter does not move the others. Because each name is keyed separately rather than drawn from one sequential stream, editing your space perturbs the parameter you edited and leaves the rest where they were.
- The generator is pinned. ChaCha8 is named explicitly, and the general-purpose "standard" and "small" generators are forbidden in the core precisely because their algorithms carry no cross-version stability guarantee — a dependency bump would silently change every study.
Sorted iteration is part of the same discipline. A trial's parameters live in a map ordered by name, spaces keep declaration order in a list, and nothing that reaches a decision, a file or an assertion iterates a hash map.
The same answer in two languages¶
The strongest form of the claim is the one that crosses the binding boundary, and it is gated.
Each paired example ends by printing a canonical block: the study name, the seed, the trial count, the sampler, and the best trial's number, value and parameters at full precision. Here is the printer, in both languages:
/// Prints the `atune-parity/1` block: the last thing this program writes, and
/// the only part `cargo dev check-parity` reads.
///
/// One `key: value` per line, one space after the colon, the fixed keys first
/// and the `best.param.*` lines last — **sorted here, by name**, rather than
/// left to the iteration order of whatever map holds the parameters. The
/// block's contract is the printer's job.
///
/// Floats print with `{:?}`, which is shortest-round-trip: the text parses back
/// to the same bits on the other side. A precision-limited format like `{:.6}`
/// would throw the comparison away, which is why the human-readable summary
/// above the block is *not* what the gate reads.
fn print_parity_block(best: &FrozenTrial, value: f64) {
println!("atune-parity/1");
println!("study: {STUDY}");
println!("seed: {SEED}");
println!("trials: {TRIALS}");
println!("sampler: {SAMPLER}");
println!("best.number: {}", best.number.get());
println!("best.value: {value:?}");
let mut params: Vec<(&String, &ParamValue)> = best.params.iter().collect();
params.sort_by_key(|(name, _)| *name);
for (name, param) in params {
println!("best.param.{name}: {}", parity_value(*param));
}
}
/// One parameter value, spelled so that Python's `repr` of the same value is
/// byte-identical.
fn parity_value(value: ParamValue) -> String {
match value {
// `{}` on an `f64` prints `1` where Python's `repr` prints `1.0`, and
// both are shortest-round-trip; `{:?}` is the spelling the two
// languages share. This example's space is float-only, so the other
// kinds are here for completeness rather than exercised by the gate.
ParamValue::F64(v) => format!("{v:?}"),
other => other.to_string(),
}
}
def print_parity_block(best: atune.FrozenTrial, value: float) -> None:
"""Prints the ``atune-parity/1`` block.
The last thing this program writes, and the only part ``cargo dev
check-parity`` reads. One ``key: value`` per line, one space after the colon,
the fixed keys first and the ``best.param.*`` lines last — **sorted here, by
name**, rather than left to the iteration order of whatever map holds the
parameters. The block's contract is the printer's job.
Floats print with ``repr``, which is shortest-round-trip: the text parses
back to the same bits on the other side. A precision-limited format like
``:.6f`` would throw the comparison away, which is why the human-readable
summary above the block is *not* what the gate reads.
"""
print("atune-parity/1")
print(f"study: {STUDY}")
print(f"seed: {SEED}")
print(f"trials: {TRIALS}")
print(f"sampler: {SAMPLER}")
print(f"best.number: {best.number}")
print(f"best.value: {value!r}")
params = best.params
for name in sorted(params):
print(f"best.param.{name}: {params[name]!r}")
cargo dev check-parity runs both arms and compares the two blocks field by
field. The header line must be byte-equal, so a format change is a version bump
rather than a silent reinterpretation. Keys must arrive in sorted name order —
required of the printer, so that the gate does not quietly depend on which
container happens to hold the parameters. String fields are compared as bytes;
float fields are parsed and compared as bits.
That last choice is the point. Rust and Python format the same f64
differently, so a text diff would fail on formatting while saying nothing about
the arithmetic. Comparing bits asserts exactly the claim being made: the two
languages computed the same number. For the Rastrigin pair that number is
best.value: 1.8650553405295156 — identical on both sides, and both arms print
with a shortest-round-trip format so the text recovers the exact bits on parse.
Seven pairs are gated today: rastrigin, pruning, multiseed, pbt,
multiobjective, open_range and resume. open_range compares a study whose
upper bound grows while it runs — the growth trajectory is inside the compared
block, so the two languages must widen at the same trials to the same bounds —
and resume compares a study reopened from a file in both languages, so the
reopen path is upstream of every compared field.
Two examples are unpaired, and the manifest records why for each:
custom_sampler is Rust by nature — a Python object cannot implement the
sampler trait, so there is nothing to compare against and never will be — and
migrate.py is an Optuna-import walkthrough whose input is a Python-made
journal, so it has no Rust twin either.
Two preconditions, both found by running the gate¶
- No fused multiply-add, and the same operation order in both arms. Rust's fused multiply-add rounds once where a separate multiply and add round twice, and Python had no counterpart before 3.13. Where the two spellings agree, that is a property of the inputs, not of the expression. The rule is enforced by the manifest's loader rather than left as advice, because the comparison itself cannot see it: with the fused spelling restored, the Rastrigin pair still passes.
- Both arms run on the same machine. Cosine, logarithm and their neighbours defer to the platform's C math library, which is not required to be correctly rounded and may differ between platforms. Compared on one machine, the gate is a statement about atune; compared across two, it would be a statement about libm. The committed goldens record the platform they were blessed on for the same reason.
Where determinism stops¶
The contract above is about stateless samplers. Everything else in atune is replayable rather than pre-determined — its recorded history explains it completely, but it is not computable from the seed in advance. The distinction is worth keeping, and every affected component states its own standing:
| What | Standing |
|---|---|
History-dependent samplers (Tpe, Dehb, Nsga2, Cmaes, GpEi, Carbs) |
Single worker with a fixed seed: fully reproducible. Several workers: replayable — which trials have finished when a decision is taken is thread scheduling. |
| Every real scheduler | The same, for the same reason: it decides from the trials it can see. |
| Enqueued trials | The values are exactly what you enqueued; which trial number they land on is decided by which worker claims them. |
| Retried trials | Two runs agree on the multiset of configurations retried, not necessarily on which number carries which. |
| A wall-clock budget | How many trials fit in a deadline depends on the machine. |
| An objective using ambient entropy | Nothing can save it. Take randomness from the seed the trial hands you. |
An open search space (open, around) |
Single worker with a fixed seed: fully reproducible. Several workers: replayable — an ask-time view is a sparse prefix, so a bound can move at a different trial. The recorded per-trial range may be narrower than a later dense replay computes (recorded ⊆ replayed); the record is authoritative, the replay canonical — two correct answers to different questions. |
| Float bits across platforms | Bit-exactness holds within a platform, not across. The remedy is per-platform goldens, never a tolerance — and the band test behind an expansion decision crosses ln for a continuous float, so the same-machine precondition extends to growth decisions (integer and stepped decisions, and every default-factor bound, are exact). |
This is why every example on this site that uses a pruner, a population or a model-based sampler runs one worker. It is also why a parity example that could not be made deterministic would be excluded from the gate with a written reason, rather than quietly left out.
Writing an objective that keeps the promise¶
The library hands your objective the seeds it should use — one for the objective's own randomness, and one per replicate inside a multi-seed fan. Use them, and derive any noise arithmetically from them rather than from a language random-number generator: two languages' generators disagree, and a process-global generator that nothing seeds is not reproducible in one language either.
That holds for a program atune spawns exactly as it holds for a closure it
calls, and it is the case most easily got wrong, because the seam is the
environment rather than an argument list. A tuned script is handed its per-trial
objective seed as ATUNE_SEED, next to
the trial number the study
derived it from. A script that reaches for the clock instead has opted out of
everything on this page, and nothing downstream can tell that it did — the study
will look perfectly reproducible and will not be.
Do that, in either shape, and a trial is reproducible from the study seed and its number alone — the property everything else on this page is built on.
Where to go next¶
| If you want to | Go to |
|---|---|
| See the promise in a first program | Your first study |
| Know which sampler is stateless | Samplers and schedulers |
| Know why the trial number is the key | Studies and trials |
| Stop and resume without changing the answer | Storage, Resume and scale |
| Seed a program atune spawns, rather than a closure it calls | Environment variables |
| Use seeds as an evaluation protocol, not just a knob | Multi-seed RL protocol |