Skip to content

Benchmarks

This page publishes measurements. It contains no verdict, no recommendation and no "faster than" claim: every table is preceded by the command that produces it, the parameters that command ran with, and the versions recorded by that artifact for the tools involved — and where atune's arm lost, the loss is in the same table as the wins.

Two things follow from that rule, and both are deliberate. Numbers that cannot be reproduced from this repository are not here at all. And numbers this repository can produce but has not measured with a spread are marked as such rather than dressed up.

The suite also writes a self-contained interactive page — every chart inline, no external asset — and a site deployment publishes it alongside these tables: the interactive benchmark page. It is the same crates/atune_bench/bench/benchmark.html that is committed in the repository, so a clone opens it without a browser ever touching this site.

What is measured, and what is not

The large suite measures samplers on synthetic single-objective problems, driven through kurobako — Optuna's own benchmark harness — so that atune's arms and an Optuna reference arm see the same problems, the same budget and the same seeds through the same scheduler.

It does not measure the multi-fidelity path: no pruning, no scheduler and no reinforcement-learning workload appears in it. It also does not measure the Python or CLI surfaces; both delegate to the same core, and that the Rust and Python arms agree bit-for-bit is asserted by the parity gate rather than benchmarked.

Two measurement arms live outside this page, as tests that gate with the suite rather than as published tables: crates/atune_bench/tests/open_range.rs compares a fixed range, the manual diagnose-and-widen loop and the automatic growth policy at a matched budget (primary metric: trials to reach a threshold), and crates/atune_bench/tests/calibration.rs measures the margins of the open-space detection constants. Their current numbers are printed by the tests themselves and recorded in the feature's spec.

No RL benchmark has been run in this repository. The reasoning behind the RL-oriented features, and the published literature it rests on, is on RL evidence — which is argument and citation, not measurement.

The kurobako suite

Provenance, as recorded by the committed summary rather than by a flag: the native kurobako version is unavailable in this historical artifact; problems kurobako_problems=0.1.14, and solver arms stamped atune_bench=0.1.0 (six arms), kurobako_solvers=0.2.2 (one arm) and optuna=4.9.0, kurobako-py=0.2.1 (four arms). 11 arms, 13 problems, budget 100 trials per study, 20 repeats, seeds 20260725–20260744, 2840 studies. The artifacts stamp their own run date and measured commit (the run: line in RESULTS.md and summary.json's provenance; a +dirty suffix would mean the worktree carried uncommitted changes when the suite ran). The committed artifacts are from the 2026-08-26 run at the clean 9f783fa — re-run twice that day, the second time purely to stamp a clean tree, with every value cell byte-identical between the two runs.

The tables below are committed-artifact evidence, not a fresh checkout measurement. Command 3 writes a candidate artifact but does not update this page. The benchmark workflow may run the same script on scheduled or manual jobs; its output remains a candidate until a reviewed change replaces the committed artifact.

Reproduce the recipe — three commands, in order:

Use a clean checkout at the artifact's recorded commit, 9f783fa634f68c3d01c6ac223e62410abafb6a17. Before running the commands, verify that git status --porcelain is empty and that git rev-parse HEAD prints that commit. The Python command consumes the checked-in universal hash lock; --require-hashes rejects an artifact not listed there, so do not replace it with an unpinned install. That lock was added after the 2026-08-26 run: it makes a future candidate refresh's dependency closure exact, but cannot recover package bytes or tool versions absent from the historical summary. Run the commands from a clean checkout of the revision that contains this lock; the historical checkout is provenance evidence and predates the lock.

  1. cargo install kurobako --version 0.2.10 --locked
  2. python3 -m venv .venv && .venv/bin/pip install --require-hashes --only-binary=:all: -r .github/release/python-benchmark-requirements.txt
  3. BUDGET=100 REPEATS=20 SEED=20260725 PYTHON=$PWD/.venv/bin/python ARTIFACTS=crates/atune_bench/bench crates/atune_bench/bench/run_suite.sh ~/kurobako-out

Every number in this section is quoted from the committed artifacts (summary.json is the machine-readable source for the generated aggregate and Ackley tables below; RESULTS.md carries the same rendered values). The recipe pins a future candidate refresh, but the historical summary did not record the full dependency closure or host/toolchain versions, so it is not a claim that a fresh run will reproduce the committed cells byte-for-byte.

The committed summary also predates the full environment record. Its generated headers therefore say (unavailable) for metadata it does not contain; the 0.2.10 harness pin above is the documented reproduction recipe, not a value recovered from summary.json.

Cross-problem aggregate

Over the 11 problems every arm ran; every cell below is from command 3 above. normalized is 0 for the best arm in a cell and 1 for the worst; mean rank averages tied arms rather than ordering them by name; win rates are paired by seed, with a tie counting as half.

Arm normalized (0 = best) mean rank win rate vs kurobako-random win rate vs optuna-tpe note
atune-gp 0.032 1.55 100% 94%
atune-auto 0.076 1.95 100% 89%
optuna-cmaes 0.236 4.64 90% 63%
atune-tpe 0.250 4.50 93% 59%
atune-cmaes 0.265 4.73 88% 59%
optuna-tpe 0.267 4.73 90% —
optuna-qmc 0.674 8.00 54% 11%
atune-sobol 0.692 8.55 48% 11% deterministic (ignores the seed)
kurobako-random 0.700 9.00 — 10%
optuna-random 0.729 9.09 45% 8%
atune-random 0.730 9.27 49% 8%

Two problems are excluded from that aggregate, and both exclusions are structural rather than chosen after seeing the result:

  • ln(sigopt/evalset/Ackley(dim=2)) — kurobako's ln wrapper evaluates the same objective at ln x, so it is protocol coverage rather than a second landscape. Measured agreement: 214 of 220 (arm, seed) cells reached an identical final value.
  • sigopt/evalset/Sphere(dim=6, int=[0, 1, 2, 3, 4, 5]) — atune-cmaes declines a space with no continuous axis, so scoring the problem would average the other ten arms over a subset one arm never ran.

One problem in full

sigopt/evalset/Ackley(dim=6), final best value, lower is better, again from command 3 above. It is here because it is a problem where atune's arms lose: atune-tpe is behind optuna-tpe, atune-cmaes behind optuna-cmaes, and atune-random behind optuna-random.

Arm best (mean ± sd) repeats solver time / study
atune-auto 3.34236 ± 1.15 20 8122.0 ms
atune-gp 3.34236 ± 1.15 20 7320.8 ms
optuna-tpe 8.36072 ± 1.47 20 585.9 ms
optuna-cmaes 8.95716 ± 1.73 20 118.5 ms
atune-cmaes 8.96967 ± 1.14 20 3.7 ms
atune-tpe 10.257 ± 1.93 20 39.5 ms
optuna-random 13.9797 ± 1.55 20 55.9 ms
atune-random 14.5717 ± 1.74 20 2.7 ms
kurobako-random 15.2465 ± 1.51 20 0.0 ms
atune-sobol 15.6586 ± 0 20 2.1 ms
optuna-qmc 15.6586 ± 0 20 89.6 ms

The other twelve per-problem tables are in the same artifact. Four readings the table does not make for you:

  • A ± 0 means unseeded, not reliable. An unscrambled Sobol sequence ignores the seed, so those 20 repeats are one sample shown 20 times.
  • Solver time is reported, never ranked on. The Optuna arms pay a Python round-trip per protocol message that the native arms do not, and the two Gaussian-process arms pay a model fit per suggestion.
  • The win-rate baseline is kurobako-random, which shares no code with atune, rather than atune's own random arm.
  • Optuna's GPSampler is not an arm. It needs torch, which the benchmark environment does not carry, so atune's GP arm has no direct counterpart here.

The in-repo sampler oracle

A second, much smaller measurement runs inside the ordinary test gate: a self-contained oracle over standard problems and three samplers, with no external harness and nothing to install.

Reproduce it with one command: cargo run --locked --release -p atune_bench --bin atune-bench -- 120 8

Budget 120 trials, mean of 8 seeds. This tool prints no dispersion, so every cell below is a mean of 8 runs and not a distribution; the suite above is where per-repeat spread is reported. The cells trace to a committed artifact: with ATUNE_BENCH_JSON=<path> the same run writes its numbers as JSON, and crates/atune_bench/bench/oracle.json is that file for the tables below. The oracle does not capture a Rust toolchain or host record, so its exact bytes are evidence for the committed source and lockfile, not a claim about an independently recovered compiler or machine. With those inputs fixed, the sweep is deterministic and re-running the command reproduces it exactly. Final gap to the known optimum, lower is better:

Problem random sobol tpe
sphere, 2-D 3.0121e-1 0.0000e0 5.9029e-4
sphere, 5-D 5.8348e0 0.0000e0 4.9250e-1
shifted sphere, 2-D 3.3623e-1 8.2421e-2 3.7727e-4
shifted sphere, 5-D 8.8523e0 9.3999e0 6.9139e-1
Rastrigin, 2-D 5.4198e0 0.0000e0 4.0237e-1
Ackley, 2-D 9.0886e0 0.0000e0 9.9977e-1
Branin 3.4350e-1 3.0654e-1 5.6728e-2

Sobol's 0.0000e0 cells are a property of the fixture rather than a result: those four problems are centred on the origin, and an unscrambled Sobol sequence evaluates the domain centre early, so the sampler lands on the optimum exactly before any search has happened. The shifted variants exist to move the optimum off that grid — and there Sobol loses to random search in 5-D (bold above), while on Branin its margin is 3.6962e-2. That is why the crate's own quality gate asserts the Sobol ordering only where it holds.

The same command ends with a multi-fidelity block: budget 1440 fidelity units, mean of 8 seeds, quality measured as the true value of the best configuration re-evaluated at the full budget.

Problem dehb random tpe
sphere 4-D, ladder [1, 3, 9] 4.1093e-1 (400 evaluations) 2.8038e0 (160) 1.1051e-1 (160)
Rastrigin 3-D, ladder [1, 3, 9] 3.7679e0 (400) 1.4250e1 (160) 2.2134e0 (160)
sphere 4-D, ladder [1, 3, 9, 27] 1.1144e0 (252) 6.5478e0 (54) 7.5370e-1 (54)

DEHB loses to full-budget TPE on all three of these surfaces. The other number in the same cells is a count rather than a quality: at a matched fidelity budget DEHB evaluated two to five times as many configurations as the other two arms. Both are stated; neither is a verdict.

How this page stays honest

  • A number without a runnable command is not published. Several measurements this project has made are therefore absent — every result that depends on an external agent implementation among them.
  • The artifact is regenerated; the narrative stays hand-authored. The marked aggregate and Ackley tables are rendered from committed summary.json (stamped with its run date and commit), and the oracle tables from the committed bench/oracle.json the same command writes under ATUNE_BENCH_JSON. A re-run refreshes the artifacts, and this page is edited to follow them.
  • A correction is visible. A number found to be wrong is corrected with a note saying so, not silently replaced.

Where to go next

If you want to Go to
The reasoning behind the RL-oriented features RL evidence
Which tool has which capability, with citations Prior art
What atune does that is unusual Why atune
The samplers and schedulers themselves Catalog
To run these gates yourself Contributing