Skip to main content

Crate atune_bench

Crate atune_bench 

Expand description

Benchmark surface for atune: the M2.3 sampler-quality gate.

This crate is the standing oracle that guards sampler quality from the moment a non-trivial sampler exists (TPE, M2.2). It has two faces over one shared engine (harness):

  • The self-contained oracle (this library + its tests). A set of standard problems (problem) run against each built-in sampler (sampler) for a fixed budget and seed, producing best-value curves. The regression test asserts the quality ordering that must always hold — TPE and Sobol beat Random at a matched budget on smooth problems, by a margin — and pins a TPE baseline curve with a documented tolerance so a quality regression fails the build. It is a normal cargo test, so it runs in the release gate as well as the nightly benchmark job. The atune-bench binary prints the same curves for a human to read.

  • The kurobako solver adapter (kurobako). A program (atune-solver) that speaks kurobako’s JSON solver protocol on stdin/stdout, so Optuna’s kurobako benchmark harness can drive atune as a solver on a machine that has kurobako installed. It has been verified end to end against kurobako 0.2.10 (2026-07-25; see the kurobako module docs for what that run proved). CI has no kurobako binary, so the evidence is kept as data: the captured wire traffic in tests/fixtures/kurobako-0.2.10/ is replayed through the real binary by tests/kurobako_wire.rs, alongside the protocol-shape and message-cycle tests.

Both faces drive the real atune_core samplers at the sampler seam (docs/design/03-architecture.md §3); nothing here re-implements a sampler, and everything is deterministic given (sampler, problem, seed, budget) (harness).

§The scheduler oracle (M3.2)

A third surface guards the schedulers rather than the samplers. Where the sampler gate asks “does the sampler find good points”, the scheduler oracle asks “does the pruner stop the right trials” — a pruner that stops a would-have-won trial still “runs”, so merely observing that pruning happened proves nothing. synthetic fabricates reproducible learning curves (noisy, saturating, deceptive, slow-starting) and scheduler_bench runs the real Study loop over them under a given scheduler, recording which trials were pruned, at what step, and with what value. The scheduler_oracle integration test turns those records into the assertions the M3.2 gate is specified as (docs/design/09-implementation.md §12).

§The PBT oracle (M4.1)

A fourth surface guards Pbt, the first fork-capable scheduler. Where the scheduler oracle asks “does the pruner stop the right trials”, the PBT oracle asks “does population-based training discover a hyperparameter schedule a fixed hyperparameter cannot reach” (docs/design/04-rl-and-oniro.md §A.2). pbt_bench builds a deliberately nonstationary problem — the best learning rate changes over training — and drives the real Study loop with the real Pbt scheduler, so the Fork path, the checkpoint-reference hand-off and the recorded lineage are exercised end-to-end. The pbt_oracle integration test asserts that PBT beats the best fixed hyperparameter and random search at a matched budget, and that the winning lineage carries the discovered schedule (docs/design/09-implementation.md §13).

§The DEHB oracle (M4.2)

A fifth surface guards Dehb, the first multi-fidelity sampler. Its claim is different in kind from every earlier one: not “a better point per trial” but “a better point per resource unit”, earned by evaluating most configurations cheaply and paying the full budget only for survivors (docs/design/02-field-analysis.md §2.1). So dehb_bench builds a problem with a genuinely correlated-but-imperfect cheap proxy (MultiFidelity) and matches every comparison on total fidelity units (Budget::with_fidelity_units), never on trial count. The dehb_oracle integration test asserts that DEHB beats random search at a matched unit budget, that it reaches a full-fidelity-only searcher’s quality for materially fewer units, and that a DE-crippled DEHB loses the margin (docs/design/09-implementation.md §13).

Every quality claim there is a win-rate or an aggregate over a 32-seed set with thresholds taken from a measured 64-seed sweep, never a per-seed universal: a stochastic searcher does not beat another on every seed, and an oracle that says it does is pinned to the seeds it was written against.

§The CMA-ES oracle (M4.3)

A sixth surface guards Cmaes, the continuous specialist. Beating random search on a smooth bowl is table stakes — TPE has done that since M2.2 — so cmaes_bench asks the question CMA-ES exists to answer: does the adapted covariance solve a problem that is ill-conditioned and rotated off the coordinate frame, where an axis-aligned searcher (uniform random, Sobol, or a TPE whose densities are per-dimension) has no representation for the valley? Landscape builds that fixture and its two controls (a sphere; the same ellipsoid unrotated), and IsotropicEs — CMA-ES with the covariance update deleted, C ≡ I, everything else identical down to the RNG stream — is the regression sentinel that proves the margin is the covariance’s doing. The cmaes_oracle integration test states the same way: win-rates over 32 seeds, thresholds from a measured 64-seed sweep, and the honest comparison against TPE recorded whichever way it falls.

§The NSGA-II oracle (M4.4)

A seventh surface guards Nsga2, the multi-objective sampler — and it is the one whose quality is a different shape from every other oracle here. A multi-objective study has no best value (StudyView::best answers None on purpose), so there is no best-so-far Curve to compare: quality is front quality, and a front is good along three separable axes — coverage (hypervolume), convergence (distance to the analytically known true front) and diversity (is it spread, or one point cloned). nsga2_bench builds four problems with knowable answers (ZDT1, the concave ZDT2, the three-objective DTLZ2 and Deb’s constrained CONSTR) and measures all three axes; CrowdlessNsga2 — NSGA-II with crowding distance deleted, the same operators and the same RNG stream otherwise — is the regression sentinel that proves the diversity axis bites, since a hypervolume-only gate cannot see the difference. The nsga2_oracle integration test states every claim as a win-rate over 32 seeds with thresholds from a measured 64-seed sweep, and records the honest comparisons (including the ones NSGA-II loses).

§The fANOVA ranking-agreement cross-check (M5.4)

An eighth surface guards the in-house PED-ANOVA importance evaluator (D21, docs/design/09-implementation.md §8.2). Its correctness gate is not a quality curve but an independent estimator: forest-fANOVA (the fanova crate) and PED-ANOVA are different algorithms — a random-forest variance decomposition versus a closed-form Parzen-density divergence — so their importance values differ by construction, but on a study with a knowable answer they must agree on the ranking. fanova_bench builds the column-major fANOVA input from a study’s completed trials and returns the forest ranking; the fanova_oracle integration test runs both estimators on the same analytic studies (y = a·x1 + b·x2 + …, a ≫ b ≫ …) and asserts the robust agreement — both put the dominant parameter first — while printing both importance vectors so the divergence in the values is visible (§8.3: assert the top-k agreement that survives, not the noisy tail). The fanova crate is bench-only and appears nowhere on the shipped path.

§The FreezeThaw oracle (M6.1)

A ninth surface guards FreezeThaw, the first scheduler that emits Decision::Pause and Command::ResumeTrial (docs/design/09-implementation.md §15). Where the scheduler oracle asks “does the pruner stop the right trials”, this asks “does the freeze-thaw scheduler freeze the right trials and thaw the right ones” — a mechanism that is only real if the Pause/Resume seam the study loop wired at M4.0 actually executes end to end. freeze_thaw_bench drives the real Study loop over deceptive and slow-blooming learning curves, wrapping the scheduler so it records the exact set of trials paused and the exact set resumed, and honours a resume by continuing the trial from its checkpoint step. The freeze_thaw_oracle integration test asserts those exact sets (the §8.2 “exact set, not just ‘it paused’” discipline), that a paused trial ends Paused keeping its checkpoint and a resumed trial continues past its pause step, and — over a seed sweep — the honest §8.3 claim that freeze-thaw spends fewer resource-units than running every trial to completion, printing the quality margin.

§The online-tuner ablation (M6.0)

A tenth surface guards the in-loop OnlineTuner, the real-time tuner that adapts hyperparameters inside one training run (docs/design/09-implementation.md §15, D19). Its claim is a different shape again: not “a better point per trial” but “tracks a moving optimum a fixed choice cannot”. So online_bench builds a deliberately non-stationary in-loop problem — the best knob value drifts across three regimes — and drives three strategies over the same reward realizations: the clustered-UCB online tuner, the best fixed arm in hindsight (the strongest static baseline), and a PBT-lite continuous-knob tuner. The online_oracle integration test leads with the mechanical, seed-free claim (on the noise-free instance the online tuner tracks each regime’s optimum and so beats the best fixed arm), then over a 64-seed sweep asserts the robust §8.3 win-rate of online over static — with headroom below the measured value — and prints the full ablation table (online vs static vs pbt-lite), the D19 in-repo artifact.

Re-exports§

pub use cmaes_bench::CONDITION;
pub use cmaes_bench::Landscape;
pub use cmaes_bench::RunReport;
pub use cmaes_bench::Searcher;
pub use cmaes_bench::rayleigh;
pub use cmaes_bench::run as run_landscape;
pub use cmaes_bench::run_cmaes;
pub use crowdless_nsga2::CrowdlessNsga2;
pub use dehb_bench::DeKnobs;
pub use dehb_bench::MfReport;
pub use dehb_bench::MultiFidelity;
pub use dehb_bench::run_dehb;
pub use dehb_bench::run_full_fidelity;
pub use dehb_bench::run_on_ladder;
pub use fanova_bench::FOREST_SEED;
pub use fanova_bench::fanova_ranking;
pub use harness::optimize;
pub use isotropic_es::IsotropicEs;
pub use nsga2_bench::MoProblem;
pub use nsga2_bench::MoReport;
pub use nsga2_bench::MoSearcher;
pub use nsga2_bench::POPULATION;
pub use nsga2_bench::run as run_multi_objective;
pub use nsga2_bench::run_with_population as run_multi_objective_with_population;
pub use online_bench::ARMS;
pub use online_bench::AblationRow;
pub use online_bench::N_ROUNDS;
pub use online_bench::OnlineReport;
pub use online_bench::REGIME_LEN;
pub use online_bench::ablation_row;
pub use online_bench::optimal_arm;
pub use online_bench::reward;
pub use online_bench::run_online;
pub use online_bench::run_pbt_lite;
pub use online_bench::run_static_best;
pub use pbt_bench::Member;
pub use pbt_bench::PbtReport;
pub use pbt_bench::ScheduleProblem;
pub use pbt_bench::random_search_best;
pub use pbt_bench::run_pbt;
pub use pbt_bench::tuned_pbt;
pub use problem::CostProblem;
pub use problem::Curve;
pub use problem::Problem;
pub use sampler::SamplerKind;
pub use scheduler_bench::BenchReport;
pub use scheduler_bench::TrialOutcome;
pub use scheduler_bench::run as run_scheduler;
pub use synthetic::LabeledCurve;
pub use synthetic::LearningCurve;

Modules§

cmaes_bench
The M4.3 CMA-ES oracle’s driver: prove Cmaes earns its slot as the continuous black-box optimizer — and prove it where the covariance is what does the earning (docs/design/09-implementation.md §13).
crowdless_nsga2
The regression sentinel for Nsga2: the same genetic algorithm with crowding distance deleted.
dehb_bench
The M4.2 DEHB oracle’s driver: prove Dehb finds a better configuration than random search per resource unit spent, not per trial.
fanova_bench
Forest-fANOVA importance, for cross-checking the in-house PED-ANOVA ranking.
freeze_thaw_bench
The FreezeThaw oracle’s driver: run a real Study with FreezeThaw over synthetic curves and record the exact set of trials it pauses and the exact set it resumes.
harness
The sampler-driving engine, at the seam and nothing above it.
isotropic_es
The regression sentinel for Cmaes: the same evolution strategy with the covariance update deleted.
kurobako
The kurobako solver protocol, implemented against atune’s samplers.
nsga2_bench
The M4.4 NSGA-II oracle’s driver: prove Nsga2 finds a Pareto front, not a point (docs/design/09-implementation.md §13).
online_bench
The online-tuner ablation’s driver: a synthetic non-stationary in-loop tuning problem, and the three strategies compared on it (M6.0, docs/design/09-implementation.md §15, §8.3; docs/design/04-rl-and-oniro.md §A.3).
pb2_bench
The M6.3 PB2 oracle’s driver: prove Pb2 — the time-varying GP-bandit variant of PBT — beats a random-perturbation PBT on a non-stationary problem where modelling (time, hyperparameter) → improvement genuinely pays (docs/design/09-implementation.md §8.3, §15).
pbt_bench
The M4.1 PBT oracle’s driver: prove Pbt discovers a hyperparameter schedule on a nonstationary problem and beats the best fixed hyperparameter and random search at a matched budget.
problem
Standard black-box optimization problems and the best-value curve.
sampler
The samplers the benchmark surface exposes behind a flag.
scheduler_bench
The scheduler oracle’s driver: run a real Study with a scheduler over synthetic curves and record exactly what it pruned.
synthetic
Synthetic learning curves — the raw material of the scheduler oracle.