Skip to main content

Module kurobako

Module kurobako 

Expand description

The kurobako solver protocol, implemented against atune’s samplers.

kurobako is Optuna’s Rust benchmark harness. It drives an external solver over a pair of pipes, one newline-delimited JSON message per line (kurobako_core’s epi — external program interface). This module is the solver side of that protocol: kurobako ask becomes an atune sampler suggestion, kurobako tell becomes an atune observation.

§The message flow, exactly

  1. On startup the solver emits one SolverMessage::SolverSpecCast telling kurobako its name and the search-space kinds it can handle.
  2. Then it serves a request loop, keyed by solver_id so one process can back several concurrent studies:
    • CreateSolverCast — a cast (no reply): translate the problem’s parameter domain into an atune SpaceSchema and start a session seeded from kurobako’s random_seed.
    • AskCall — reply AskReply: sample the next point, encode it positionally as Vec<f64> in the domain’s variable order.
    • TellCall — reply TellReply: fold the evaluated objective back into the session’s history and run the sampler’s after_trial callback (see Learning below).
    • DropSolverCast — drop the session.

§Learning: this adapter is a driver, not a seam oracle

crate::harness drives the Sampler seam and deliberately stops at suggest/record, because a seam oracle should measure the seam. This module is not that: it is a driver, standing where Study stands in a real run, so it owes the sampler the same callbacks the real loop makes.

The load-bearing one is after_trial. It is the only channel through which a stateful sampler learns: Cmaes releases the trial’s in-flight offspring there and folds its loss into the weighted recombination, so a driver that skips it runs a CMA-ES whose covariance never adapts — a sampler that looks alive (it answers every ask, in range, deterministically) and scores at uniform-random level. The real loop’s ordering is mirrored exactly (atune_core::study::handle’s finish): the trial is folded into the view first, so the snapshot the sampler sees already contains the finished trial, and the callback then runs on that view plus the finished trial itself. Every terminal state gets the call — Complete and Failed — because the real loop calls it from tell_complete, tell_pruned and tell_failed alike, and because a stateful sampler needs the failure to release the pending entry it recorded at ask time. A failure reported by the callback becomes an ErrorReply rather than being swallowed.

§Trial ids

kurobako owns the trial-id counter and sends its next value as next_trial_id on every AskCall; the solver assigns that id to the new trial and replies with the counter advanced by one (it consumes exactly one id per ask). Internally the session numbers trials densely from zero — that internal number is what seeds the sampler — and maps kurobako’s id onto it so a later tell finds the right trial.

§Domain translation

kurobako range + distributionatune Distribution
CONTINUOUS + UNIFORMFloat { low, high, log: false }
CONTINUOUS + LOG_UNIFORMFloat { low, high, log: true }
DISCRETE [low, high) + UNIFORMInt(low, high − 1)
DISCRETE [low, high) + LOG_UNIFORMInt(low, high − 1, log)
CATEGORICALCat(choices)

kurobako’s numeric ranges are half-open (high exclusive); atune’s are closed. For a discrete range the exclusive upper bound maps to the inclusive high − 1. For a continuous range the single excluded endpoint has measure zero, so it is mapped directly — the only faithful choice given closed atune ranges, and documented here rather than hidden.

Three of the five rows are pinned by captured kurobako wire traffic (see Verified against kurobako 0.2.10 below): CONTINUOUS+UNIFORM, CONTINUOUS+LOG_UNIFORM and DISCRETE+UNIFORM. The other two have no fixture, because no built-in problem reaches them; each passed one bounded live check through a temporary external problem, and is otherwise held by the in-tree tests:

  • CATEGORICAL — no kurobako built-in problem offers a categorical domain without a multi-gigabyte dataset (nasbench, hpobench).
  • DISCRETE + LOG_UNIFORM — no built-in emits it. The obvious candidate is the ln wrapper, but it only converts continuous variables: kurobako problem ln "$(kurobako problem sigopt --dim 3 --int=0 --int=1 --int=2 sphere)" still transmits DISCRETE+UNIFORM (checked against kurobako 0.2.10). The row exists because the wire format allows it — Range::Discrete and LOG_UNIFORM are independent fields, so a third-party problem can pair them and would otherwise hit an unreachable translation.

§Fidelity is not modelled

These samplers do not prune or do multi-fidelity, so every ask requests evaluation to the problem’s last step (next_step = last), exactly as kurobako’s own random solver does in its non-multi-fidelity mode. The advertised Capability therefore omit Concurrent, MultiObjective and Conditional; the internal history model already tracks in-flight (Running) trials, so enabling Concurrent later is a one-line capability change — but it stays off until a live concurrent run has actually been driven through this adapter. That reticence is what protects the numbers: kurobako refuses to start a study whose solver lacks a capability the problem needs (Error: Incapable … incapables=[MultiObjective] on a ZDT problem), so an un-implemented capability that we do not advertise costs a benchmark run, while one we advertise falsely would produce plausible, wrong numbers. What is advertised is therefore asserted by equality, not by spot checks, in the_spec_message_advertises_exactly_the_verified_capabilities.

§Verified against kurobako 0.2.10 (2026-07-25)

This adapter has been driven end to end by the real kurobako binary (kurobako 0.2.10, kurobako_problems=0.1.14). What that run established, and how each fact is now held in place without the binary:

  • It optimizes. atune-tpe reached best 2.941616 ± 0.155291 against kurobako’s own random at 9.592990 ± 2.185300 on sigopt Ackley(dim=2), budget 40 × 3 repeats. Those are seed-specific numbers — every arm of a kurobako study is reseeded from --seed, so a different seed moves both columns — and they reproduce only from exactly this invocation:

    $ SOLVER=$(kurobako solver --name atune-tpe command -- \
        "$(cargo metadata --format-version 1 --no-deps \
           | sed -n 's/.*"target_directory":"\([^"]*\)".*/\1/p')/release/atune-solver" \
        --sampler tpe)
    $ ACKLEY=$(kurobako problem sigopt --dim 2 ackley)
    $ kurobako studies --solvers "$SOLVER" "$(kurobako solver random)" \
        --problems "$ACKLEY" --budget 40 --repeats 3 --seed 7 \
      | kurobako run --quiet | kurobako report

    --seed 7 is part of the claim, not decoration: seeds 3 and 20260725 give different (still TPE-favouring) numbers.

  • Continuous UNIFORM, DISCRETE, the message flow and the framing are pinned by the captured sessions in tests/fixtures/kurobako-0.2.10/*.jsonl (the exact bytes kurobako sent), replayed through the real binary by tests/kurobako_wire.rs.

  • LOG_UNIFORM is the exact inverse of kurobako’s own log transform. kurobako’s problem ln wrapper turns a UNIFORM [a, b] domain into LOG_UNIFORM [eᵃ, eᵇ] and evaluates f(ln x). Running the same sampler and the same random_seed on both produced byte-identical objective values for every trial, with the suggested parameters differing by exactly exp (3.974182045260828 ↔ 53.20657852888938). Nothing but a correct log transform on both sides does that, and the invariant is asserted from the captured pair.

  • Multi-objective is refused by kurobako itself, because this solver does not advertise the capability (see above).

  • The stateful arm really adapts. atune-cmaes reached 0.001012 ± 0.000584 against atune-random at 0.155862 ± 0.208828 on sigopt Sphere(dim=2), budget 100 × 3 repeats, --seed 7 — a 154× gap on a convex bowl. This is the live counterpart of the Learning section above: before after_trial was wired, the same comparison came out a tie (0.132 vs 0.161 at budget 100 × 20 repeats), which is what a CMA-ES with a frozen covariance looks like from outside.

  • An exhausted search space no longer aborts the run. --sampler auto --budget 100 on sigopt Sphere(dim=2, int=[0,1]) — a 49-point domain — used to end the whole recipe with Error: Other (cause; search space exhausted) after 50 asks. It now completes, with one stderr line naming the switch at trial 49 (see A search space that runs out).

  • CATEGORICAL and DISCRETE+LOG_UNIFORM translate live. On 2026-09-19 kurobako 0.2.10 drove the TPE solver through two temporary external problems (a three-choice categorical domain and a discrete [1, 9) log-uniform one), ten evaluations each: every parameter was in domain and no ERROR_REPLY was sent. That is protocol coverage, not a performance claim, and no fixture captures it (see Domain translation).

Concurrency, conditionals and multi-step (pruning) problems remain unexercised live.

§A search space that runs out

SamplerKind::Auto can resolve to Grid — its rule 3 picks exhaustive enumeration for a finite space of at most 256 points when the declared budget covers it — and Grid reports Error::SpaceExhausted once every point has been handed out. atune’s own study loop treats that as a normal stop. kurobako’s protocol has no way to say it: a solver can only answer an ask or send an ErrorReply, and kurobako treats an error as fatal for the whole run — it aborts, throwing away every completed study of every other solver and problem in the same recipe. A three-point grid could therefore delete an hour of someone else’s benchmark.

So this adapter keeps answering. On the first SpaceExhausted the session permanently substitutes uniform Random for the rest of that study, writes exactly one explanatory line to stderr, and carries on. What that means for reading a result:

  • It is a harness accommodation, not a sampler. The atune-auto arm’s score past the exhaustion point measures uniform random search, not auto’s pick. A run whose stderr carries that line is reporting a mixture, and the trial at which it switched is the boundary.
  • The arm stays faithful to what auto_sampler actually picked. The substitution happens at exhaustion, never at construction: swapping Grid out up front would benchmark a sampler auto did not choose and quietly hide the rule that produced it.
  • It cannot move the headline number. Exhaustion means the grid has already handed out every point of a finite space, so the domain’s optimum is in the history before the first substituted draw. The substitution can therefore only re-evaluate ground already covered: Best is settled, and only a rate metric like kurobako’s AUC sees the tail at all.

Structs§

Adapter
The solver process: one sampler kind, and a session per live solver_id.
Domain
A parameter domain: an ordered list of variables (mirror of kurobako_core::domain::Domain, a newtype over the vector).
EvaluatedTrial
An evaluated trial (mirror of kurobako_core::trial::EvaluatedTrial).
NextTrial
A trial to evaluate (mirror of kurobako_core::trial::NextTrial).
NullableF64Vec
A Vec<f64> that round-trips non-finite entries as JSON null.
ProblemSpec
A problem specification — only the fields the solver reads.
SolverSpec
The solver’s specification (mirror of kurobako_core::solver::SolverSpec).
Variable
One domain variable (mirror of kurobako_core::domain::Variable).

Enums§

Capability
A solver capability (mirror of kurobako_core::solver::Capability).
ErrorKind
A kurobako error kind (mirror of kurobako_core::ErrorKind).
KurobakoDistribution
The prior distribution of a variable (mirror of kurobako_core::domain::Distribution).
Range
A variable’s value range (mirror of kurobako_core::domain::Range).
SolverMessage
A message on the kurobako solver protocol.
Steps
The evaluable steps of a problem (mirror of kurobako_core::problem::EvaluableSteps, serialized untagged).

Constants§

DEFAULT_BUDGET_TRIALS
The trial budget an Adapter assumes when the caller does not say.