Module kurobako
Expand description
The kurobako solver protocol, implemented against atune’s samplers.
kurobako is Optuna’s Rust benchmark
harness. It drives an external solver over a pair of pipes, one
newline-delimited JSON message per line (kurobako_core’s epi — external
program interface). This module is the solver side of that protocol:
kurobako ask becomes an atune sampler suggestion, kurobako tell becomes
an atune observation.
§The message flow, exactly
- On startup the solver emits one
SolverMessage::SolverSpecCasttelling kurobako its name and the search-space kinds it can handle. - Then it serves a request loop, keyed by
solver_idso one process can back several concurrent studies:CreateSolverCast— a cast (no reply): translate the problem’s parameter domain into an atuneSpaceSchemaand start a session seeded from kurobako’srandom_seed.AskCall— replyAskReply: sample the next point, encode it positionally asVec<f64>in the domain’s variable order.TellCall— replyTellReply: fold the evaluated objective back into the session’s history and run the sampler’safter_trialcallback (see Learning below).DropSolverCast— drop the session.
§Learning: this adapter is a driver, not a seam oracle
crate::harness drives the Sampler seam
and deliberately stops at suggest/record, because a seam oracle should
measure the seam. This module is not that: it is a driver, standing where
Study stands in a real run, so it owes the
sampler the same callbacks the real loop makes.
The load-bearing one is
after_trial. It is the only
channel through which a stateful sampler learns: Cmaes
releases the trial’s in-flight offspring there and folds its loss into the
weighted recombination, so a driver that skips it runs a CMA-ES whose
covariance never adapts — a sampler that looks alive (it answers every ask,
in range, deterministically) and scores at uniform-random level. The real
loop’s ordering is mirrored exactly (atune_core::study::handle’s finish):
the trial is folded into the view first, so the snapshot the sampler sees
already contains the finished trial, and the callback then runs on that
view plus the finished trial itself. Every terminal state gets the call —
Complete and Failed — because the real loop calls it from
tell_complete, tell_pruned and tell_failed alike, and because a
stateful sampler needs the failure to release the pending entry it recorded
at ask time. A failure reported by the callback becomes an
ErrorReply rather than being swallowed.
§Trial ids
kurobako owns the trial-id counter and sends its next value as
next_trial_id on every AskCall; the solver assigns that id to the new
trial and replies with the counter advanced by one (it consumes exactly one
id per ask). Internally the session numbers trials densely from zero — that
internal number is what seeds the sampler — and maps kurobako’s id onto it so
a later tell finds the right trial.
§Domain translation
| kurobako range + distribution | atune Distribution |
|---|---|
CONTINUOUS + UNIFORM | Float { low, high, log: false } |
CONTINUOUS + LOG_UNIFORM | Float { low, high, log: true } |
DISCRETE [low, high) + UNIFORM | Int(low, high − 1) |
DISCRETE [low, high) + LOG_UNIFORM | Int(low, high − 1, log) |
CATEGORICAL | Cat(choices) |
kurobako’s numeric ranges are half-open (high exclusive); atune’s are
closed. For a discrete range the exclusive upper bound maps to the inclusive
high − 1. For a continuous range the single excluded endpoint has measure
zero, so it is mapped directly — the only faithful choice given closed atune
ranges, and documented here rather than hidden.
Three of the five rows are pinned by captured kurobako wire traffic (see
Verified against kurobako 0.2.10 below): CONTINUOUS+UNIFORM,
CONTINUOUS+LOG_UNIFORM and DISCRETE+UNIFORM. The other two have no
fixture, because no built-in problem reaches them; each passed one bounded
live check through a temporary external problem, and is otherwise held by
the in-tree tests:
CATEGORICAL— no kurobako built-in problem offers a categorical domain without a multi-gigabyte dataset (nasbench,hpobench).DISCRETE+LOG_UNIFORM— no built-in emits it. The obvious candidate is thelnwrapper, but it only converts continuous variables:kurobako problem ln "$(kurobako problem sigopt --dim 3 --int=0 --int=1 --int=2 sphere)"still transmitsDISCRETE+UNIFORM(checked against kurobako 0.2.10). The row exists because the wire format allows it —Range::DiscreteandLOG_UNIFORMare independent fields, so a third-party problem can pair them and would otherwise hit an unreachable translation.
§Fidelity is not modelled
These samplers do not prune or do multi-fidelity, so every ask requests
evaluation to the problem’s last step (next_step = last), exactly as
kurobako’s own random solver does in its non-multi-fidelity mode. The
advertised Capability therefore omit Concurrent, MultiObjective
and Conditional; the internal history model already tracks in-flight
(Running) trials, so enabling Concurrent later is a one-line capability
change — but it stays off until a live concurrent run has actually been
driven through this adapter. That reticence is what protects the numbers:
kurobako refuses to start a study whose solver lacks a capability the
problem needs (Error: Incapable … incapables=[MultiObjective] on a ZDT
problem), so an un-implemented capability that we do not advertise costs a
benchmark run, while one we advertise falsely would produce plausible,
wrong numbers. What is advertised is therefore asserted by equality, not by
spot checks, in
the_spec_message_advertises_exactly_the_verified_capabilities.
§Verified against kurobako 0.2.10 (2026-07-25)
This adapter has been driven end to end by the real kurobako binary
(kurobako 0.2.10, kurobako_problems=0.1.14). What that run established,
and how each fact is now held in place without the binary:
-
It optimizes.
atune-tpereached best2.941616 ± 0.155291against kurobako’s ownrandomat9.592990 ± 2.185300onsigoptAckley(dim=2), budget 40 × 3 repeats. Those are seed-specific numbers — every arm of a kurobako study is reseeded from--seed, so a different seed moves both columns — and they reproduce only from exactly this invocation:$ SOLVER=$(kurobako solver --name atune-tpe command -- \ "$(cargo metadata --format-version 1 --no-deps \ | sed -n 's/.*"target_directory":"\([^"]*\)".*/\1/p')/release/atune-solver" \ --sampler tpe) $ ACKLEY=$(kurobako problem sigopt --dim 2 ackley) $ kurobako studies --solvers "$SOLVER" "$(kurobako solver random)" \ --problems "$ACKLEY" --budget 40 --repeats 3 --seed 7 \ | kurobako run --quiet | kurobako report--seed 7is part of the claim, not decoration: seeds 3 and 20260725 give different (still TPE-favouring) numbers. -
Continuous
UNIFORM,DISCRETE, the message flow and the framing are pinned by the captured sessions intests/fixtures/kurobako-0.2.10/*.jsonl(the exact bytes kurobako sent), replayed through the real binary bytests/kurobako_wire.rs. -
LOG_UNIFORMis the exact inverse of kurobako’s own log transform. kurobako’sproblem lnwrapper turns aUNIFORM [a, b]domain intoLOG_UNIFORM [eᵃ, eᵇ]and evaluatesf(ln x). Running the same sampler and the samerandom_seedon both produced byte-identical objective values for every trial, with the suggested parameters differing by exactlyexp(3.974182045260828↔53.20657852888938). Nothing but a correct log transform on both sides does that, and the invariant is asserted from the captured pair. -
Multi-objective is refused by kurobako itself, because this solver does not advertise the capability (see above).
-
The stateful arm really adapts.
atune-cmaesreached0.001012 ± 0.000584againstatune-randomat0.155862 ± 0.208828onsigoptSphere(dim=2), budget 100 × 3 repeats,--seed 7— a 154× gap on a convex bowl. This is the live counterpart of the Learning section above: beforeafter_trialwas wired, the same comparison came out a tie (0.132vs0.161at budget 100 × 20 repeats), which is what a CMA-ES with a frozen covariance looks like from outside. -
An exhausted search space no longer aborts the run.
--sampler auto --budget 100onsigoptSphere(dim=2, int=[0,1])— a 49-point domain — used to end the whole recipe withError: Other (cause; search space exhausted)after 50 asks. It now completes, with one stderr line naming the switch at trial 49 (see A search space that runs out). -
CATEGORICALandDISCRETE+LOG_UNIFORMtranslate live. On 2026-09-19 kurobako 0.2.10 drove the TPE solver through two temporary external problems (a three-choice categorical domain and a discrete[1, 9)log-uniform one), ten evaluations each: every parameter was in domain and noERROR_REPLYwas sent. That is protocol coverage, not a performance claim, and no fixture captures it (see Domain translation).
Concurrency, conditionals and multi-step (pruning) problems remain unexercised live.
§A search space that runs out
SamplerKind::Auto can resolve to
Grid — its rule 3 picks exhaustive enumeration
for a finite space of at most 256 points when the declared budget covers it —
and Grid reports Error::SpaceExhausted
once every point has been handed out. atune’s own study loop treats that as a
normal stop. kurobako’s protocol has no way to say it: a solver can only
answer an ask or send an ErrorReply, and
kurobako treats an error as fatal for the whole run — it aborts, throwing
away every completed study of every other solver and problem in the same
recipe. A three-point grid could therefore delete an hour of someone else’s
benchmark.
So this adapter keeps answering. On the first SpaceExhausted the session
permanently substitutes uniform Random
for the rest of that study, writes exactly one explanatory line to stderr,
and carries on. What that means for reading a result:
- It is a harness accommodation, not a sampler. The
atune-autoarm’s score past the exhaustion point measures uniform random search, notauto’s pick. A run whose stderr carries that line is reporting a mixture, and the trial at which it switched is the boundary. - The arm stays faithful to what
auto_sampleractually picked. The substitution happens at exhaustion, never at construction: swappingGridout up front would benchmark a samplerautodid not choose and quietly hide the rule that produced it. - It cannot move the headline number. Exhaustion means the grid has
already handed out every point of a finite space, so the domain’s optimum
is in the history before the first substituted draw. The substitution can
therefore only re-evaluate ground already covered:
Bestis settled, and only a rate metric like kurobako’sAUCsees the tail at all.
Structs§
- Adapter
- The solver process: one sampler kind, and a session per live
solver_id. - Domain
- A parameter domain: an ordered list of variables (mirror of
kurobako_core::domain::Domain, a newtype over the vector). - Evaluated
Trial - An evaluated trial (mirror of
kurobako_core::trial::EvaluatedTrial). - Next
Trial - A trial to evaluate (mirror of
kurobako_core::trial::NextTrial). - Nullable
F64Vec - A
Vec<f64>that round-trips non-finite entries as JSONnull. - Problem
Spec - A problem specification — only the fields the solver reads.
- Solver
Spec - The solver’s specification (mirror of
kurobako_core::solver::SolverSpec). - Variable
- One domain variable (mirror of
kurobako_core::domain::Variable).
Enums§
- Capability
- A solver capability (mirror of
kurobako_core::solver::Capability). - Error
Kind - A kurobako error kind (mirror of
kurobako_core::ErrorKind). - Kurobako
Distribution - The prior distribution of a variable (mirror of
kurobako_core::domain::Distribution). - Range
- A variable’s value range (mirror of
kurobako_core::domain::Range). - Solver
Message - A message on the kurobako solver protocol.
- Steps
- The evaluable steps of a problem (mirror of
kurobako_core::problem::EvaluableSteps, serialized untagged).
Constants§
- DEFAULT_
BUDGET_ TRIALS - The trial budget an
Adapterassumes when the caller does not say.