Skip to main content

Module online

Module online 

Expand description

In-loop online tuning — a real-time hyperparameter tuner that lives inside one training run.

This is the headline RL feature only a native tuner can offer: the ask/tell Study loop tunes across runs, but an OnlineTuner adapts hyperparameters during a single run, from the reward signal it is fed each round. Its decision cost is bounded and tiny — the suggest/observe hot path is O(clusters), allocation-free and CPU-pure (no linalg, no GPU, D4) — so it is safe to call on oniro’s real-time side (async SAC at 500 Hz with hard deadlines).

use atune_core::online::{Mabc, OnlineTuner};

let mut tuner = OnlineTuner::builder()
    .cluster("lr", &[1e-4, 3e-4, 1e-3])
    .cluster("ent_coef", &[0.0, 0.01, 0.02])
    .policy(Mabc::ucb(32)) // ULTHO-style clustered UCB, window = 32
    .build()?;

for _round in 0..3 {
    let (lr, ent) = {
        let hp = tuner.suggest(); // zero-alloc, O(clusters)
        (hp.get("lr").unwrap(), hp.get("ent_coef").unwrap())
    };
    // ...train one round with `lr` / `ent`, producing a return...
    let round_return = 1.0 - (lr - 3e-4).abs() - ent; // stand-in
    tuner.observe(round_return); // sliding-window credit
}

§The two policies

  • Mabc — the ULTHO-style clustered sliding-window UCB bandit over discrete candidate values per cluster. The primary policy; the exact UCB rule is pinned and golden-tested.
  • PbtLite — a deterministic scalar-perturbation policy for continuous knobs: periodic ×up/×down explore steps, no forking, no RNG.

§Determinism and portability

Both policies are deterministic by construction — the UCB bandit needs no RNG at all, and PBT-lite’s perturbation direction is a pure function of the reward stream — so a run is reproducible from its inputs alone (the determinism contract). The module is atune_core, so it inherits the wasm floor and the ground rules: no faer, no std::time, no threads, and no allocation on the hot path.

§Still a replayable study

The hot path touches no storage, but the tuner can be snapshotted into a StateBlob at interval boundaries — never on the hot path — so an online run remains a replayable, resumable study (snapshot/restore, is_flush_boundary).

§The oniro seam

oniro consumes the tuner through a per-round callback (the oniro PR1 shape) — atune stays oniro-agnostic, so that integration is author-gated. drive_round is the oniro-free shape of that callback: suggest → run one round → observe. The honest ablation lives in atune_bench (online_oracle, §8.3).

Structs§

Choice
The set of hyperparameter choices OnlineTuner::suggest produced this round — one value per cluster, borrowed from the tuner (no allocation).
Mabc
The ULTHO-style clustered UCB policy: a sliding-window UCB1 bandit over each cluster’s discrete candidate values.
OnlineTuner
A real-time, in-loop hyperparameter tuner.
OnlineTunerBuilder
A builder for an OnlineTuner.
PbtLite
The deterministic PBT-lite scalar-perturbation policy for continuous knobs.

Enums§

Policy
Which in-loop policy an OnlineTuner runs. Constructed from a Mabc or a PbtLite via OnlineTunerBuilder::policy (both impl Into<Policy>).

Constants§

DEFAULT_EXPLORATION
The default UCB exploration constant c (the classic UCB1 value √2).
DEFAULT_PERTURB_DOWN
The default PBT-lite explore down factor.
DEFAULT_PERTURB_UP
The default PBT-lite explore up factor.
DEFAULT_WINDOW
The default sliding window for Mabc::ucb, in rounds.
MAX_WINDOW
The largest Mabc window accepted — bounds the ring-buffer allocation so a user-supplied window can never size an unbounded allocation.
ONLINE_BLOB_KIND
The kind tag of the persisted OnlineTuner snapshot (StateBlob).
ONLINE_BLOB_VERSION
The schema version of the ONLINE_BLOB_KIND payload.