Module online
Expand description
In-loop online tuning — a real-time hyperparameter tuner that lives inside one training run.
This is the headline RL feature only a native tuner can offer: the ask/tell
Study loop tunes across runs, but an
OnlineTuner adapts hyperparameters during a single run, from the reward
signal it is fed each round. Its decision cost is bounded and tiny — the
suggest/observe hot path is
O(clusters), allocation-free and CPU-pure (no linalg, no GPU, D4) — so
it is safe to call on oniro’s real-time side (async SAC at 500 Hz with hard
deadlines).
use atune_core::online::{Mabc, OnlineTuner};
let mut tuner = OnlineTuner::builder()
.cluster("lr", &[1e-4, 3e-4, 1e-3])
.cluster("ent_coef", &[0.0, 0.01, 0.02])
.policy(Mabc::ucb(32)) // ULTHO-style clustered UCB, window = 32
.build()?;
for _round in 0..3 {
let (lr, ent) = {
let hp = tuner.suggest(); // zero-alloc, O(clusters)
(hp.get("lr").unwrap(), hp.get("ent_coef").unwrap())
};
// ...train one round with `lr` / `ent`, producing a return...
let round_return = 1.0 - (lr - 3e-4).abs() - ent; // stand-in
tuner.observe(round_return); // sliding-window credit
}§The two policies
Mabc— the ULTHO-style clustered sliding-window UCB bandit over discrete candidate values per cluster. The primary policy; the exact UCB rule is pinned and golden-tested.PbtLite— a deterministic scalar-perturbation policy for continuous knobs: periodic×up/×downexplore steps, no forking, no RNG.
§Determinism and portability
Both policies are deterministic by construction — the UCB bandit needs no
RNG at all, and PBT-lite’s perturbation direction is a pure function of the
reward stream — so a run is reproducible from its inputs alone (the
determinism
contract).
The module is atune_core, so it inherits the wasm floor and the ground
rules: no faer, no std::time, no threads, and no allocation on the hot
path.
§Still a replayable study
The hot path touches no storage, but the tuner can be snapshotted into a
StateBlob at interval boundaries — never on the
hot path — so an online run remains a replayable, resumable study
(snapshot/restore,
is_flush_boundary).
§The oniro seam
oniro consumes the tuner through a per-round callback (the oniro PR1
shape) — atune stays oniro-agnostic, so
that integration is author-gated. drive_round is
the oniro-free shape of that callback: suggest → run one round → observe. The
honest ablation lives in atune_bench (online_oracle, §8.3).
Structs§
- Choice
- The set of hyperparameter choices
OnlineTuner::suggestproduced this round — one value per cluster, borrowed from the tuner (no allocation). - Mabc
- The ULTHO-style clustered UCB policy: a sliding-window UCB1 bandit over each cluster’s discrete candidate values.
- Online
Tuner - A real-time, in-loop hyperparameter tuner.
- Online
Tuner Builder - A builder for an
OnlineTuner. - PbtLite
- The deterministic PBT-lite scalar-perturbation policy for continuous knobs.
Enums§
- Policy
- Which in-loop policy an
OnlineTunerruns. Constructed from aMabcor aPbtLiteviaOnlineTunerBuilder::policy(bothimpl Into<Policy>).
Constants§
- DEFAULT_
EXPLORATION - The default UCB exploration constant
c(the classic UCB1 value√2). - DEFAULT_
PERTURB_ DOWN - The default PBT-lite explore down factor.
- DEFAULT_
PERTURB_ UP - The default PBT-lite explore up factor.
- DEFAULT_
WINDOW - The default sliding window for
Mabc::ucb, in rounds. - MAX_
WINDOW - The largest
Mabcwindow accepted — bounds the ring-buffer allocation so a user-supplied window can never size an unbounded allocation. - ONLINE_
BLOB_ KIND - The
kindtag of the persistedOnlineTunersnapshot (StateBlob). - ONLINE_
BLOB_ VERSION - The schema version of the
ONLINE_BLOB_KINDpayload.