Skip to main content

Module dehb_bench

Module dehb_bench 

Expand description

The M4.2 DEHB oracle’s driver: prove Dehb finds a better configuration than random search per resource unit spent, not per trial.

This is to DEHB what pbt_bench is to PBT and scheduler_bench is to the pruners (docs/design/09-implementation.md §13). DEHB is a multi-fidelity optimizer: its claim is not “better points per trial” — it is “better points per env-step”, earned by evaluating most configurations at a cheap low fidelity and spending the expensive full budget only on the survivors (docs/design/02-field-analysis.md §2.1, docs/design/03-architecture.md §3.1). So every comparison here is at a matched total resource budget (Budget::with_fidelity_units), never at a matched trial count — at a matched trial count DEHB would silently be handed max/min times the compute (9× on the default ladder, 27× on the wide one), which would prove nothing.

§The multi-fidelity problem

MultiFidelity wraps one of the standard synthetic surfaces (Problem::sphere, Problem::rastrigin) as its full-fidelity truth and adds the two things that make a problem multi-fidelity at all:

  • a correlated cheap proxy. A configuration evaluated at fidelity b reads as f(x) + decay(b)·(bias(x) + wobble), where bias(x) is a smooth bowl centred off the true optimum. A cheap evaluation therefore ranks configurations correlated with — but not identical to — the truth, which is exactly the property multi-fidelity search exploits and the property a too-easy fixture would fake. rank_correlation measures it, and the oracle pins it into a band so the difficulty of the fixture cannot silently drift.
  • a learning-curve decay: decay(b) = (b_max/b − 1) / (b_max/b_min − 1), so the proxy’s error is full at b_min, has mostly burned off by the time the ladder reaches b_max/η, and is exactly zero at b_max. The full-budget evaluation is the ground truth, which is what makes the incumbent rule below fair to every searcher.

§How quality is measured (the same rule for every searcher)

The incumbent is the best configuration a run has evaluated at the maximum fidelity, and its quality is that configuration’s true value. Both halves matter:

  • a run may only claim a configuration it actually paid full budget for — so a lucky low-fidelity reading can never become the headline number;
  • and because the full-budget observation is exact, “best observed at b_max” is “best true among the fully-evaluated”, with no selection on information the searcher did not have.

MfReport::curve records that incumbent against cumulative units, so the two claims the oracle asserts are both readable off one run: quality-at-matched-budget (MfReport::incumbent_true) and budget-to-matched-quality (MfReport::units_to_reach).

§How DEHB drives the fidelity through the real loop

Each atune trial is one DEHB job: the sampler breeds a configuration for a target fidelity and the objective runs it to exactly that budget — Dehb::target_fidelity is the hand-off, and ctx.report(fidelity, …) both records the value and charges the study’s fidelity budget (the loop charges a trial its highest reported step). Successive halving lives inside DEHB (its promotion queue), so no pruner is paired here: the multi-fidelity gating is the schedule itself, and pairing a HyperbandPruner over the same ladder would be a second, redundant realization of it.

§The baselines

BaselineWhat it isolates
run_full_fidelity with SamplerKind::Randomplain random search — the claim “DEHB beats random at a matched unit budget”
run_full_fidelity with SamplerKind::Tpea strong full-fidelity-only searcher — it is what keeps the claim honest about where DEHB does not win
run_on_ladderthe identical DEHB fidelity schedule with random configurations instead of DE-bred ones — isolates what differential evolution adds on top of the multi-fidelity schedule
run_dehb with DeKnobs::degenerateDEHB whose DE is crippled (F = 0, CR = 0: the child is its own target) — the regression sentinel

§Why DEHB is not a SamplerKind

The M2.3 sampler gate (harness, SamplerKind, tests/regression.rs) compares samplers at a matched trial count. Adding Dehb there would hand it max_fidelity / min_fidelity times the compute of every other sampler and call the result a win — the exact mistake this module exists to avoid. Dehb therefore has its own surface, on its own axis, and the trial-count gate keeps guarding the single-fidelity samplers only.

§Determinism

Every run is single-worker (parallelism(1)) with a fixed study seed, a ManualClock, and a pure objective (the proxy wobble is a seeded hash of the configuration, the fidelity and the seed), so a whole run — trial order, fidelity schedule, incumbent and unit curve — is reproducible. DEHB is history-dependent and stateful, so it gets single-worker determinism plus replayability, never trial-number-indexed determinism (docs/design/09-implementation.md §5); the oracle pins the single-worker property and says nothing stronger.

Structs§

DeKnobs
The DE knobs a run is driven with.
MfReport
What one search run achieved, in resource units.
MfTrial
One evaluated trial, as the oracle reads it back.
MultiFidelity
A standard synthetic surface plus a cheap, correlated low-fidelity proxy.

Constants§

ETA
The reduction factor every fixture uses — Hyperband’s standard η = 3.

Functions§

fidelity_schedule
The fidelity DEHB assigns to each of its first jobs jobs.
run_dehb
Runs DEHB over problem until the fidelity budget is spent.
run_full_fidelity
Runs a full-fidelity-only search: every trial pays the full budget.
run_on_ladder
Runs a searcher on DEHB’s own fidelity schedule — the DE ablation.