Benchmarks¶
This page publishes measurements. It contains no verdict, no recommendation and no "faster than" claim: every table is preceded by the command that produces it, the parameters that command ran with, and the versions recorded by that artifact for the tools involved — and where atune's arm lost, the loss is in the same table as the wins.
Two things follow from that rule, and both are deliberate. Numbers that cannot be reproduced from this repository are not here at all. And numbers this repository can produce but has not measured with a spread are marked as such rather than dressed up.
The suite also writes a self-contained interactive page — every chart inline, no
external asset — and a site deployment publishes it alongside these tables:
the interactive benchmark page. It is the same
crates/atune_bench/bench/benchmark.html that is committed in the repository, so
a clone opens it without a browser ever touching this site.
What is measured, and what is not¶
The large suite measures samplers on synthetic single-objective problems, driven through kurobako — Optuna's own benchmark harness — so that atune's arms and an Optuna reference arm see the same problems, the same budget and the same seeds through the same scheduler.
It does not measure the multi-fidelity path: no pruning, no scheduler and no reinforcement-learning workload appears in it. It also does not measure the Python or CLI surfaces; both delegate to the same core, and that the Rust and Python arms agree bit-for-bit is asserted by the parity gate rather than benchmarked.
Two measurement arms live outside this page, as tests that gate with the
suite rather than as published tables: crates/atune_bench/tests/open_range.rs
compares a fixed range, the manual diagnose-and-widen loop and the automatic
growth policy at a matched budget (primary metric: trials to reach a
threshold), and crates/atune_bench/tests/calibration.rs measures the
margins of the open-space
detection constants. Their current numbers are printed by the tests
themselves and recorded in the feature's spec.
No RL benchmark has been run in this repository. The reasoning behind the RL-oriented features, and the published literature it rests on, is on RL evidence — which is argument and citation, not measurement.
The kurobako suite¶
Provenance, as recorded by the committed summary rather than by a flag:
the native kurobako version is unavailable in this historical artifact;
problems kurobako_problems=0.1.14, and solver arms stamped
atune_bench=0.1.0 (six arms), kurobako_solvers=0.2.2 (one arm) and
optuna=4.9.0, kurobako-py=0.2.1 (four arms). 11 arms, 13 problems, budget 100
trials per study, 20 repeats, seeds 20260725–20260744, 2840 studies. The
artifacts stamp their own run date and measured commit (the run: line in
RESULTS.md and summary.json's provenance; a +dirty suffix would mean
the worktree carried uncommitted changes when the suite ran). The committed
artifacts are from the 2026-08-26 run at the clean 9f783fa — re-run twice
that day, the second time purely to stamp a clean tree, with every value cell
byte-identical between the two runs.
The tables below are committed-artifact evidence, not a fresh checkout measurement. Command 3 writes a candidate artifact but does not update this page. The benchmark workflow may run the same script on scheduled or manual jobs; its output remains a candidate until a reviewed change replaces the committed artifact.
Reproduce the recipe — three commands, in order:
Use a clean checkout at the artifact's recorded commit, 9f783fa634f68c3d01c6ac223e62410abafb6a17.
Before running the commands, verify that git status --porcelain is empty and
that git rev-parse HEAD prints that commit. The Python command consumes the
checked-in universal hash lock; --require-hashes rejects an artifact not
listed there, so do not replace it with an unpinned install. That lock was added
after the 2026-08-26 run: it makes a future candidate refresh's dependency
closure exact, but cannot recover package bytes or tool versions absent from the
historical summary. Run the commands from a clean checkout of the revision that
contains this lock; the historical checkout is provenance evidence and predates
the lock.
cargo install kurobako --version 0.2.10 --lockedpython3 -m venv .venv && .venv/bin/pip install --require-hashes --only-binary=:all: -r .github/release/python-benchmark-requirements.txtBUDGET=100 REPEATS=20 SEED=20260725 PYTHON=$PWD/.venv/bin/python ARTIFACTS=crates/atune_bench/bench crates/atune_bench/bench/run_suite.sh ~/kurobako-out
Every number in this section is quoted from the committed artifacts
(summary.json is the machine-readable source for the generated aggregate and
Ackley tables below; RESULTS.md carries the same rendered values). The recipe
pins a future candidate refresh, but the historical summary did not record the
full dependency closure or host/toolchain versions, so it is not a claim that a
fresh run will reproduce the committed cells byte-for-byte.
The committed summary also predates the full environment record. Its generated
headers therefore say (unavailable) for metadata it does not contain; the
0.2.10 harness pin above is the documented reproduction recipe, not a value
recovered from summary.json.
Cross-problem aggregate¶
Over the 11 problems every arm ran; every cell below is from command 3 above.
normalized is 0 for the best arm in a cell and 1 for the worst; mean rank
averages tied arms rather than ordering them by name; win rates are paired by
seed, with a tie counting as half.
| Arm | normalized (0 = best) | mean rank | win rate vs kurobako-random |
win rate vs optuna-tpe |
note |
|---|---|---|---|---|---|
atune-gp |
0.032 | 1.55 | 100% | 94% | |
atune-auto |
0.076 | 1.95 | 100% | 89% | |
optuna-cmaes |
0.236 | 4.64 | 90% | 63% | |
atune-tpe |
0.250 | 4.50 | 93% | 59% | |
atune-cmaes |
0.265 | 4.73 | 88% | 59% | |
optuna-tpe |
0.267 | 4.73 | 90% | — | |
optuna-qmc |
0.674 | 8.00 | 54% | 11% | |
atune-sobol |
0.692 | 8.55 | 48% | 11% | deterministic (ignores the seed) |
kurobako-random |
0.700 | 9.00 | — | 10% | |
optuna-random |
0.729 | 9.09 | 45% | 8% | |
atune-random |
0.730 | 9.27 | 49% | 8% |
Two problems are excluded from that aggregate, and both exclusions are structural rather than chosen after seeing the result:
ln(sigopt/evalset/Ackley(dim=2))— kurobako'slnwrapper evaluates the same objective atln x, so it is protocol coverage rather than a second landscape. Measured agreement: 214 of 220 (arm, seed) cells reached an identical final value.sigopt/evalset/Sphere(dim=6, int=[0, 1, 2, 3, 4, 5])—atune-cmaesdeclines a space with no continuous axis, so scoring the problem would average the other ten arms over a subset one arm never ran.
One problem in full¶
sigopt/evalset/Ackley(dim=6), final best value, lower is better, again from
command 3 above. It is here because it is a problem where atune's arms lose:
atune-tpe is behind optuna-tpe, atune-cmaes behind optuna-cmaes, and
atune-random behind optuna-random.
| Arm | best (mean ± sd) | repeats | solver time / study |
|---|---|---|---|
atune-auto |
3.34236 ± 1.15 | 20 | 8122.0 ms |
atune-gp |
3.34236 ± 1.15 | 20 | 7320.8 ms |
optuna-tpe |
8.36072 ± 1.47 | 20 | 585.9 ms |
optuna-cmaes |
8.95716 ± 1.73 | 20 | 118.5 ms |
atune-cmaes |
8.96967 ± 1.14 | 20 | 3.7 ms |
atune-tpe |
10.257 ± 1.93 | 20 | 39.5 ms |
optuna-random |
13.9797 ± 1.55 | 20 | 55.9 ms |
atune-random |
14.5717 ± 1.74 | 20 | 2.7 ms |
kurobako-random |
15.2465 ± 1.51 | 20 | 0.0 ms |
atune-sobol |
15.6586 ± 0 | 20 | 2.1 ms |
optuna-qmc |
15.6586 ± 0 | 20 | 89.6 ms |
The other twelve per-problem tables are in the same artifact. Four readings the table does not make for you:
- A
± 0means unseeded, not reliable. An unscrambled Sobol sequence ignores the seed, so those 20 repeats are one sample shown 20 times. - Solver time is reported, never ranked on. The Optuna arms pay a Python round-trip per protocol message that the native arms do not, and the two Gaussian-process arms pay a model fit per suggestion.
- The win-rate baseline is
kurobako-random, which shares no code with atune, rather than atune's own random arm. - Optuna's
GPSampleris not an arm. It needstorch, which the benchmark environment does not carry, so atune's GP arm has no direct counterpart here.
The in-repo sampler oracle¶
A second, much smaller measurement runs inside the ordinary test gate: a self-contained oracle over standard problems and three samplers, with no external harness and nothing to install.
Reproduce it with one command:
cargo run --locked --release -p atune_bench --bin atune-bench -- 120 8
Budget 120 trials, mean of 8 seeds. This tool prints no dispersion, so every
cell below is a mean of 8 runs and not a distribution; the suite above is where
per-repeat spread is reported. The cells trace to a committed artifact: with
ATUNE_BENCH_JSON=<path> the same run writes its numbers as JSON, and
crates/atune_bench/bench/oracle.json is that file for the tables below. The
oracle does not capture a Rust toolchain or host record, so its exact bytes are
evidence for the committed source and lockfile, not a claim about an
independently recovered compiler or machine. With those inputs fixed, the sweep
is deterministic and re-running the command reproduces it exactly.
Final gap to the known optimum, lower is better:
| Problem | random |
sobol |
tpe |
|---|---|---|---|
| sphere, 2-D | 3.0121e-1 | 0.0000e0 | 5.9029e-4 |
| sphere, 5-D | 5.8348e0 | 0.0000e0 | 4.9250e-1 |
| shifted sphere, 2-D | 3.3623e-1 | 8.2421e-2 | 3.7727e-4 |
| shifted sphere, 5-D | 8.8523e0 | 9.3999e0 | 6.9139e-1 |
| Rastrigin, 2-D | 5.4198e0 | 0.0000e0 | 4.0237e-1 |
| Ackley, 2-D | 9.0886e0 | 0.0000e0 | 9.9977e-1 |
| Branin | 3.4350e-1 | 3.0654e-1 | 5.6728e-2 |
Sobol's 0.0000e0 cells are a property of the fixture rather than a result:
those four problems are centred on the origin, and an unscrambled Sobol sequence
evaluates the domain centre early, so the sampler lands on the optimum exactly
before any search has happened. The shifted variants exist to move the optimum
off that grid — and there Sobol loses to random search in 5-D (bold above),
while on Branin its margin is 3.6962e-2. That is why the crate's own quality
gate asserts the Sobol ordering only where it holds.
The same command ends with a multi-fidelity block: budget 1440 fidelity units, mean of 8 seeds, quality measured as the true value of the best configuration re-evaluated at the full budget.
| Problem | dehb |
random |
tpe |
|---|---|---|---|
| sphere 4-D, ladder [1, 3, 9] | 4.1093e-1 (400 evaluations) | 2.8038e0 (160) | 1.1051e-1 (160) |
| Rastrigin 3-D, ladder [1, 3, 9] | 3.7679e0 (400) | 1.4250e1 (160) | 2.2134e0 (160) |
| sphere 4-D, ladder [1, 3, 9, 27] | 1.1144e0 (252) | 6.5478e0 (54) | 7.5370e-1 (54) |
DEHB loses to full-budget TPE on all three of these surfaces. The other number in the same cells is a count rather than a quality: at a matched fidelity budget DEHB evaluated two to five times as many configurations as the other two arms. Both are stated; neither is a verdict.
How this page stays honest¶
- A number without a runnable command is not published. Several measurements this project has made are therefore absent — every result that depends on an external agent implementation among them.
- The artifact is regenerated; the narrative stays hand-authored. The
marked aggregate and Ackley tables are rendered from committed
summary.json(stamped with its run date and commit), and the oracle tables from the committedbench/oracle.jsonthe same command writes underATUNE_BENCH_JSON. A re-run refreshes the artifacts, and this page is edited to follow them. - A correction is visible. A number found to be wrong is corrected with a note saying so, not silently replaced.
Where to go next¶
| If you want to | Go to |
|---|---|
| The reasoning behind the RL-oriented features | RL evidence |
| Which tool has which capability, with citations | Prior art |
| What atune does that is unusual | Why atune |
| The samplers and schedulers themselves | Catalog |
| To run these gates yourself | Contributing |