6 atune sampler arms against optuna-cmaes, optuna-qmc, optuna-random, optuna-tpe
(Optuna 4.9.0), run through Optuna's own benchmark harness
(kurobako (unavailable)) on 13 problems ×
20 repeats at a budget of 100 trials. Lower is better everywhere;
every arm of a given (problem, repeat) sees the same seed, so the comparisons are paired. Run 2026-08-26 at commit 9f783fa.
Over the 11 problems every arm ran. normalized is 0 for the
best arm in a cell and 1 for the worst; mean rank averages tied arms rather than
ordering them by name; win rates are paired by seed, a tie counting as half.
| Arm | normalized (0 = best) | mean rank | win rate vs kurobako-random | win rate vs optuna-tpe |
|
|---|---|---|---|---|---|
atune-gp | 0.032 | 1.55 | 100% | 94% | |
atune-auto | 0.076 | 1.95 | 100% | 89% | |
optuna-cmaes | 0.236 | 4.64 | 90% | 63% | |
atune-tpe | 0.250 | 4.50 | 93% | 59% | |
atune-cmaes | 0.265 | 4.73 | 88% | 59% | |
optuna-tpe | 0.267 | 4.73 | 90% | — | |
optuna-qmc | 0.674 | 8.00 | 54% | 11% | |
atune-sobol | 0.692 | 8.55 | 48% | 11% | deterministic — ignores the seed |
kurobako-random | 0.700 | 9.00 | — | 10% | |
optuna-random | 0.729 | 9.09 | 45% | 8% | |
atune-random | 0.730 | 9.27 | 49% | 8% |
ln(sigopt/evalset/Ackley(dim=2)) — a log reparameterization of sigopt/evalset/Ackley(dim=2) — kurobako's `ln` wrapper evaluates the same objective at `ln x`, so it is protocol coverage rather than an independent problem (measured: 214/220 (arm, seed) cells reached an identical final value)sigopt/evalset/Sphere(dim=6, int=[0, 1, 2, 3, 4, 5]) — not run by atune-cmaesatune-autoatune-cmaesatune-gpatune-randomatune-sobolatune-tpekurobako-randomoptuna-cmaesoptuna-qmcoptuna-randomoptuna-tpemean of 20 repeats — reparameterization, not in the aggregate
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 0.77013 ± 0.58 |
atune-gp | 0.77013 ± 0.58 |
atune-cmaes | 1.2846 ± 0.86 |
atune-tpe | 1.3176 ± 1.1 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 0.8214 ± 0.62 |
atune-gp | 0.8214 ± 0.62 |
atune-cmaes | 1.2846 ± 0.86 |
atune-tpe | 1.3176 ± 1.1 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 3.3424 ± 1.1 |
atune-gp | 3.3424 ± 1.1 |
optuna-tpe | 8.3607 ± 1.5 |
optuna-cmaes | 8.9572 ± 1.7 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 5.623 ± 0.28 |
atune-gp | 5.623 ± 0.28 |
atune-tpe | 7.2021 ± 0.92 |
optuna-tpe | 7.2199 ± 1 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | -3.0991 ± 0.074 |
atune-gp | -3.0991 ± 0.074 |
optuna-tpe | -2.9344 ± 0.18 |
atune-cmaes | -2.9316 ± 0.22 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 29.923 ± 7.8 |
atune-gp | 29.923 ± 7.8 |
optuna-cmaes | 45.074 ± 7.6 |
atune-tpe | 46.416 ± 13 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-gp | 4.4857 ± 0.52 |
optuna-cmaes | 5.5906 ± 0.25 |
atune-cmaes | 5.6499 ± 0.32 |
optuna-tpe | 5.8751 ± 0.35 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | -1.0316 ± 1.8e-05 |
atune-gp | -1.0316 ± 1.8e-05 |
atune-tpe | -1.0313 ± 0.00025 |
atune-cmaes | -1.0297 ± 0.0031 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 2.6139e-05 ± 1.4e-05 |
atune-gp | 2.6139e-05 ± 1.4e-05 |
atune-cmaes | 0.0011094 ± 0.0018 |
atune-sobol | 0.0014783 ± 0 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 0 ± 0 |
atune-gp | 0 ± 0 |
optuna-cmaes | 1 ± 0.65 |
optuna-tpe | 1.05 ± 0.83 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 7.7005e-05 ± 2.7e-05 |
atune-gp | 7.7005e-05 ± 2.7e-05 |
optuna-cmaes | 1.9603 ± 0.64 |
atune-cmaes | 2.2885 ± 1.1 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-tpe | -145.1 ± 8.6 |
atune-auto | -143.72 ± 5.7 |
atune-gp | -143.72 ± 5.7 |
optuna-tpe | -138.83 ± 11 |
mean of 20 repeats
| top arms | best (mean ± sd) |
|---|---|
atune-auto | 25.234 ± 0.25 |
atune-gp | 25.234 ± 0.25 |
optuna-tpe | 25.653 ± 0.54 |
atune-tpe | 25.731 ± 0.63 |
TPESampler(multivariate=False, n_startup_trials=10)). atune's TPE is a different TPE — same startup count and candidate count, different quantile rule — so atune-tpe vs optuna-tpe compares two implementations of an algorithm family, not one algorithm in two languages.GPSampler — the natural counterpart to atune-gp — is absent because it requires torch, which this benchmark environment does not carry. Read atune-gp against optuna-tpe and optuna-cmaes.atune-auto is told the trial budget (--budget) because kurobako never transmits it and that arm's rule ladder is budget-dependent. No other arm receives it, and Optuna 4.9 has no equivalent auto-selector.Generated by crates/atune_bench/bench/summarize.py; run seed
20260725–20260744, 20 repeats. Reproduce with
crates/atune_bench/bench/run_suite.sh.