Crate atune_bench
Expand description
Benchmark surface for atune: the M2.3 sampler-quality gate.
This crate is the standing oracle that guards sampler quality from the moment
a non-trivial sampler exists (TPE, M2.2). It has two faces over one shared
engine (harness):
-
The self-contained oracle (this library + its tests). A set of standard problems (
problem) run against each built-in sampler (sampler) for a fixed budget and seed, producing best-value curves. The regression test asserts the quality ordering that must always hold — TPE and Sobol beat Random at a matched budget on smooth problems, by a margin — and pins a TPE baseline curve with a documented tolerance so a quality regression fails the build. It is a normalcargo test, so it runs in the release gate as well as the nightly benchmark job. Theatune-benchbinary prints the same curves for a human to read. -
The kurobako solver adapter (
kurobako). A program (atune-solver) that speaks kurobako’s JSON solver protocol on stdin/stdout, so Optuna’skurobakobenchmark harness can drive atune as a solver on a machine that has kurobako installed. It has been verified end to end against kurobako 0.2.10 (2026-07-25; see thekurobakomodule docs for what that run proved). CI has no kurobako binary, so the evidence is kept as data: the captured wire traffic intests/fixtures/kurobako-0.2.10/is replayed through the real binary bytests/kurobako_wire.rs, alongside the protocol-shape and message-cycle tests.
Both faces drive the real atune_core samplers at the sampler seam
(docs/design/03-architecture.md §3); nothing here re-implements a sampler, and
everything is deterministic given (sampler, problem, seed, budget)
(harness).
§The scheduler oracle (M3.2)
A third surface guards the schedulers rather than the samplers. Where the
sampler gate asks “does the sampler find good points”, the scheduler oracle
asks “does the pruner stop the right trials” — a pruner that stops a
would-have-won trial still “runs”, so merely observing that pruning happened
proves nothing. synthetic fabricates reproducible learning curves (noisy,
saturating, deceptive, slow-starting) and scheduler_bench runs the real
Study loop over them under a given scheduler,
recording which trials were pruned, at what step, and with what value. The
scheduler_oracle integration test turns those records into the assertions
the M3.2 gate is specified as (docs/design/09-implementation.md §12).
§The PBT oracle (M4.1)
A fourth surface guards Pbt, the first
fork-capable scheduler. Where the scheduler oracle asks “does the pruner stop
the right trials”, the PBT oracle asks “does population-based training
discover a hyperparameter schedule a fixed hyperparameter cannot reach”
(docs/design/04-rl-and-oniro.md §A.2). pbt_bench builds a deliberately
nonstationary problem — the best learning rate changes over training — and
drives the real Study loop with the real Pbt
scheduler, so the Fork path, the
checkpoint-reference hand-off and the recorded lineage are exercised
end-to-end. The pbt_oracle integration test asserts that PBT beats the best
fixed hyperparameter and random search at a matched budget, and that the
winning lineage carries the discovered schedule (docs/design/09-implementation.md
§13).
§The DEHB oracle (M4.2)
A fifth surface guards Dehb, the first
multi-fidelity sampler. Its claim is different in kind from every earlier
one: not “a better point per trial” but “a better point per resource unit”,
earned by evaluating most configurations cheaply and paying the full budget
only for survivors (docs/design/02-field-analysis.md §2.1). So dehb_bench
builds a problem with a genuinely correlated-but-imperfect cheap proxy
(MultiFidelity) and matches every comparison on total fidelity units
(Budget::with_fidelity_units),
never on trial count. The dehb_oracle integration test asserts that DEHB
beats random search at a matched unit budget, that it reaches a
full-fidelity-only searcher’s quality for materially fewer units, and that a
DE-crippled DEHB loses the margin (docs/design/09-implementation.md §13).
Every quality claim there is a win-rate or an aggregate over a 32-seed set with thresholds taken from a measured 64-seed sweep, never a per-seed universal: a stochastic searcher does not beat another on every seed, and an oracle that says it does is pinned to the seeds it was written against.
§The CMA-ES oracle (M4.3)
A sixth surface guards Cmaes, the continuous
specialist. Beating random search on a smooth bowl is table stakes — TPE has
done that since M2.2 — so cmaes_bench asks the question CMA-ES exists to
answer: does the adapted covariance solve a problem that is
ill-conditioned and rotated off the coordinate frame, where an axis-aligned
searcher (uniform random, Sobol, or a TPE whose densities are per-dimension)
has no representation for the valley? Landscape builds that fixture and
its two controls (a sphere; the same ellipsoid unrotated), and
IsotropicEs — CMA-ES with the covariance update deleted, C ≡ I,
everything else identical down to the RNG stream — is the regression sentinel
that proves the margin is the covariance’s doing. The cmaes_oracle
integration test states the same way: win-rates over 32 seeds, thresholds from
a measured 64-seed sweep, and the honest comparison against TPE recorded
whichever way it falls.
§The NSGA-II oracle (M4.4)
A seventh surface guards Nsga2, the
multi-objective sampler — and it is the one whose quality is a different
shape from every other oracle here. A multi-objective study has no best value
(StudyView::best answers None on
purpose), so there is no best-so-far Curve to compare: quality is
front quality, and a front is good along three separable axes —
coverage (hypervolume), convergence (distance to the analytically known
true front) and diversity (is it spread, or one point cloned). nsga2_bench
builds four problems with knowable answers (ZDT1, the concave ZDT2, the
three-objective DTLZ2 and Deb’s constrained CONSTR) and measures all three
axes; CrowdlessNsga2 — NSGA-II with crowding distance deleted, the same
operators and the same RNG stream otherwise — is the regression sentinel that
proves the diversity axis bites, since a hypervolume-only gate cannot see the
difference. The nsga2_oracle integration test states every claim as a
win-rate over 32 seeds with thresholds from a measured 64-seed sweep, and
records the honest comparisons (including the ones NSGA-II loses).
§The fANOVA ranking-agreement cross-check (M5.4)
An eighth surface guards the in-house PED-ANOVA importance evaluator
(D21, docs/design/09-implementation.md §8.2). Its correctness gate is not a
quality curve but an independent estimator: forest-fANOVA (the fanova
crate) and PED-ANOVA are different algorithms — a random-forest variance
decomposition versus a closed-form Parzen-density divergence — so their
importance values differ by construction, but on a study with a knowable
answer they must agree on the ranking. fanova_bench builds the
column-major fANOVA input from a study’s completed trials and returns the
forest ranking; the fanova_oracle integration test runs both estimators on
the same analytic studies (y = a·x1 + b·x2 + …, a ≫ b ≫ …) and asserts
the robust agreement — both put the dominant parameter first — while
printing both importance vectors so the divergence in the values is visible
(§8.3: assert the top-k agreement that survives, not the noisy tail). The
fanova crate is bench-only and appears nowhere on the shipped path.
§The FreezeThaw oracle (M6.1)
A ninth surface guards FreezeThaw, the
first scheduler that emits Decision::Pause
and Command::ResumeTrial
(docs/design/09-implementation.md §15). Where the scheduler oracle asks “does the
pruner stop the right trials”, this asks “does the freeze-thaw scheduler
freeze the right trials and thaw the right ones” — a mechanism that is
only real if the Pause/Resume seam the study loop wired at M4.0 actually
executes end to end. freeze_thaw_bench drives the real
Study loop over deceptive and slow-blooming
learning curves, wrapping the scheduler so it records the exact set of
trials paused and the exact set resumed, and honours a resume by
continuing the trial from its checkpoint step. The freeze_thaw_oracle
integration test asserts those exact sets (the §8.2 “exact set, not just ‘it
paused’” discipline), that a paused trial ends Paused keeping its checkpoint
and a resumed trial continues past its pause step, and — over a seed sweep
— the honest §8.3 claim that freeze-thaw spends fewer resource-units than
running every trial to completion, printing the quality margin.
§The online-tuner ablation (M6.0)
A tenth surface guards the in-loop OnlineTuner,
the real-time tuner that adapts hyperparameters inside one training run
(docs/design/09-implementation.md §15, D19). Its claim is a different shape again:
not “a better point per trial” but “tracks a moving optimum a fixed choice
cannot”. So online_bench builds a deliberately non-stationary in-loop
problem — the best knob value drifts across three regimes — and drives three
strategies over the same reward realizations: the clustered-UCB online tuner,
the best fixed arm in hindsight (the strongest static baseline), and a
PBT-lite continuous-knob tuner. The online_oracle integration test leads
with the mechanical, seed-free claim (on the noise-free instance the
online tuner tracks each regime’s optimum and so beats the best fixed arm),
then over a 64-seed sweep asserts the robust §8.3 win-rate of online over
static — with headroom below the measured value — and prints the full
ablation table (online vs static vs pbt-lite), the D19 in-repo artifact.
Re-exports§
pub use cmaes_bench::CONDITION;pub use cmaes_bench::Landscape;pub use cmaes_bench::RunReport;pub use cmaes_bench::Searcher;pub use cmaes_bench::rayleigh;pub use cmaes_bench::run as run_landscape;pub use cmaes_bench::run_cmaes;pub use crowdless_nsga2::CrowdlessNsga2;pub use dehb_bench::DeKnobs;pub use dehb_bench::MfReport;pub use dehb_bench::MultiFidelity;pub use dehb_bench::run_dehb;pub use dehb_bench::run_full_fidelity;pub use dehb_bench::run_on_ladder;pub use fanova_bench::FOREST_SEED;pub use fanova_bench::fanova_ranking;pub use harness::optimize;pub use isotropic_es::IsotropicEs;pub use nsga2_bench::MoProblem;pub use nsga2_bench::MoReport;pub use nsga2_bench::MoSearcher;pub use nsga2_bench::POPULATION;pub use nsga2_bench::run as run_multi_objective;pub use nsga2_bench::run_with_population as run_multi_objective_with_population;pub use online_bench::ARMS;pub use online_bench::AblationRow;pub use online_bench::N_ROUNDS;pub use online_bench::OnlineReport;pub use online_bench::REGIME_LEN;pub use online_bench::ablation_row;pub use online_bench::optimal_arm;pub use online_bench::reward;pub use online_bench::run_online;pub use online_bench::run_pbt_lite;pub use online_bench::run_static_best;pub use pbt_bench::Member;pub use pbt_bench::PbtReport;pub use pbt_bench::ScheduleProblem;pub use pbt_bench::random_search_best;pub use pbt_bench::run_pbt;pub use pbt_bench::tuned_pbt;pub use problem::CostProblem;pub use problem::Curve;pub use problem::Problem;pub use sampler::SamplerKind;pub use scheduler_bench::BenchReport;pub use scheduler_bench::TrialOutcome;pub use scheduler_bench::run as run_scheduler;pub use synthetic::LabeledCurve;pub use synthetic::LearningCurve;
Modules§
- cmaes_
bench - The M4.3 CMA-ES oracle’s driver: prove
Cmaesearns its slot as the continuous black-box optimizer — and prove it where the covariance is what does the earning (docs/design/09-implementation.md§13). - crowdless_
nsga2 - The regression sentinel for
Nsga2: the same genetic algorithm with crowding distance deleted. - dehb_
bench - The M4.2 DEHB oracle’s driver: prove
Dehbfinds a better configuration than random search per resource unit spent, not per trial. - fanova_
bench - Forest-fANOVA importance, for cross-checking the in-house PED-ANOVA ranking.
- freeze_
thaw_ bench - The
FreezeThaworacle’s driver: run a realStudywithFreezeThawover synthetic curves and record the exact set of trials it pauses and the exact set it resumes. - harness
- The sampler-driving engine, at the seam and nothing above it.
- isotropic_
es - The regression sentinel for
Cmaes: the same evolution strategy with the covariance update deleted. - kurobako
- The kurobako solver protocol, implemented against atune’s samplers.
- nsga2_
bench - The M4.4 NSGA-II oracle’s driver: prove
Nsga2finds a Pareto front, not a point (docs/design/09-implementation.md§13). - online_
bench - The online-tuner ablation’s driver: a synthetic non-stationary in-loop
tuning problem, and the three strategies compared on it (M6.0,
docs/design/09-implementation.md§15, §8.3;docs/design/04-rl-and-oniro.md§A.3). - pb2_
bench - The M6.3 PB2 oracle’s driver: prove
Pb2— the time-varying GP-bandit variant of PBT — beats a random-perturbation PBT on a non-stationary problem where modelling(time, hyperparameter) → improvementgenuinely pays (docs/design/09-implementation.md§8.3, §15). - pbt_
bench - The M4.1 PBT oracle’s driver: prove
Pbtdiscovers a hyperparameter schedule on a nonstationary problem and beats the best fixed hyperparameter and random search at a matched budget. - problem
- Standard black-box optimization problems and the best-value curve.
- sampler
- The samplers the benchmark surface exposes behind a flag.
- scheduler_
bench - The scheduler oracle’s driver: run a real
Studywith a scheduler over synthetic curves and record exactly what it pruned. - synthetic
- Synthetic learning curves — the raw material of the scheduler oracle.