Module dehb_bench
Expand description
The M4.2 DEHB oracle’s driver: prove Dehb finds a better configuration
than random search per resource unit spent, not per trial.
This is to DEHB what pbt_bench is to PBT and
scheduler_bench is to the pruners
(docs/design/09-implementation.md §13). DEHB is a multi-fidelity optimizer: its
claim is not “better points per trial” — it is “better points per env-step”,
earned by evaluating most configurations at a cheap low fidelity and spending
the expensive full budget only on the survivors
(docs/design/02-field-analysis.md §2.1, docs/design/03-architecture.md §3.1). So every
comparison here is at a matched total resource budget
(Budget::with_fidelity_units), never at a matched trial count — at a
matched trial count DEHB would silently be handed max/min times the compute
(9× on the default ladder, 27× on the wide one), which would prove nothing.
§The multi-fidelity problem
MultiFidelity wraps one of the standard synthetic surfaces
(Problem::sphere, Problem::rastrigin) as its full-fidelity truth
and adds the two things that make a problem multi-fidelity at all:
- a correlated cheap proxy. A configuration evaluated at fidelity
breads asf(x) + decay(b)·(bias(x) + wobble), wherebias(x)is a smooth bowl centred off the true optimum. A cheap evaluation therefore ranks configurations correlated with — but not identical to — the truth, which is exactly the property multi-fidelity search exploits and the property a too-easy fixture would fake.rank_correlationmeasures it, and the oracle pins it into a band so the difficulty of the fixture cannot silently drift. - a learning-curve decay:
decay(b) = (b_max/b − 1) / (b_max/b_min − 1), so the proxy’s error is full atb_min, has mostly burned off by the time the ladder reachesb_max/η, and is exactly zero atb_max. The full-budget evaluation is the ground truth, which is what makes the incumbent rule below fair to every searcher.
§How quality is measured (the same rule for every searcher)
The incumbent is the best configuration a run has evaluated at the maximum fidelity, and its quality is that configuration’s true value. Both halves matter:
- a run may only claim a configuration it actually paid full budget for — so a lucky low-fidelity reading can never become the headline number;
- and because the full-budget observation is exact, “best observed at
b_max” is “best true among the fully-evaluated”, with no selection on information the searcher did not have.
MfReport::curve records that incumbent against cumulative units, so
the two claims the oracle asserts are both readable off one run:
quality-at-matched-budget (MfReport::incumbent_true) and
budget-to-matched-quality (MfReport::units_to_reach).
§How DEHB drives the fidelity through the real loop
Each atune trial is one DEHB job: the sampler breeds a configuration for a
target fidelity and the objective runs it to exactly that budget —
Dehb::target_fidelity is the hand-off, and ctx.report(fidelity, …) both
records the value and charges the study’s fidelity budget (the loop charges a
trial its highest reported step). Successive halving lives inside DEHB (its
promotion queue), so no pruner is paired here: the multi-fidelity gating is
the schedule itself, and pairing a
HyperbandPruner over the same
ladder would be a second, redundant realization of it.
§The baselines
| Baseline | What it isolates |
|---|---|
run_full_fidelity with SamplerKind::Random | plain random search — the claim “DEHB beats random at a matched unit budget” |
run_full_fidelity with SamplerKind::Tpe | a strong full-fidelity-only searcher — it is what keeps the claim honest about where DEHB does not win |
run_on_ladder | the identical DEHB fidelity schedule with random configurations instead of DE-bred ones — isolates what differential evolution adds on top of the multi-fidelity schedule |
run_dehb with DeKnobs::degenerate | DEHB whose DE is crippled (F = 0, CR = 0: the child is its own target) — the regression sentinel |
§Why DEHB is not a SamplerKind
The M2.3 sampler gate (harness, SamplerKind,
tests/regression.rs) compares samplers at a matched trial count. Adding
Dehb there would hand it max_fidelity / min_fidelity times the compute of
every other sampler and call the result a win — the exact mistake this module
exists to avoid. Dehb therefore has its own surface, on its own axis, and
the trial-count gate keeps guarding the single-fidelity samplers only.
§Determinism
Every run is single-worker (parallelism(1)) with a fixed study seed, a
ManualClock, and a pure objective (the proxy wobble is a seeded hash of the
configuration, the fidelity and the seed), so a whole run — trial order,
fidelity schedule, incumbent and unit curve — is reproducible. DEHB is
history-dependent and stateful, so it gets single-worker determinism plus
replayability, never trial-number-indexed determinism
(docs/design/09-implementation.md §5); the oracle pins the single-worker property
and says nothing stronger.
Structs§
- DeKnobs
- The DE knobs a run is driven with.
- MfReport
- What one search run achieved, in resource units.
- MfTrial
- One evaluated trial, as the oracle reads it back.
- Multi
Fidelity - A standard synthetic surface plus a cheap, correlated low-fidelity proxy.
Constants§
- ETA
- The reduction factor every fixture uses — Hyperband’s standard
η = 3.
Functions§
- fidelity_
schedule - The fidelity DEHB assigns to each of its first
jobsjobs. - run_
dehb - Runs DEHB over
problemuntil the fidelity budget is spent. - run_
full_ fidelity - Runs a full-fidelity-only search: every trial pays the full budget.
- run_
on_ ladder - Runs a searcher on DEHB’s own fidelity schedule — the DE ablation.