Skip to content

Resume and scale a study

Goal: survive an interrupted run, and add workers — threads, processes or machines — without a coordinator.

Both halves of this page are the same idea: a study lives in storage, not in the handle that created it. Point a study at a file and it outlives the process. Point several workers at the same storage and they are the same study. There is no scheduler daemon and no server to stand up — storage is the only thing workers coordinate through.

Give the study a file

An in-memory study dies with the process. A journal-backed one does not: it is an append-only file that records every operation, so a later handle re-reads it, continues the trial numbering, restores whatever state the sampler and scheduler persisted, and carries on.

/// Runs the study in two halves, with the handle dropped in between.
///
/// The second handle knows only the file and the study id. Everything else — the
/// seed, the direction, the declared space, the seed protocol, the trial
/// numbering, and any state the sampler or scheduler persisted — comes back out
/// of storage, which is why a resumed study cannot silently disagree with the one
/// it continues.
fn run_in_two_halves(path: &Path) -> Result<FrozenTrial> {
    let spec = path.to_string_lossy().into_owned();

    // First handle: create the study in the journal file and spend half the
    // budget. Dropping it at the end of this scope is the "interruption" — a
    // crashed worker, a pre-empted spot instance, a laptop closing.
    {
        let study = Study::builder()
            .storage(atune::open_storage(&spec)?)
            .budget(Budget::trials(SEGMENT))
            .create(StudyConfig::new(STUDY).with_seed(SEED))?;
        study.optimize(evaluate)?;
        expect_count(&study, SEGMENT, "the first handle")?;
    }

    // Second handle: a fresh `open_storage` on the same path, then `load` rather
    // than `create`. The budget counts the trials already in the study, so
    // `trials(TRIALS)` means "run up to TRIALS", not "run TRIALS more".
    let study = Study::builder()
        .storage(atune::open_storage(&spec)?)
        .budget(Budget::trials(TRIALS))
        .load(ONLY_STUDY)?;
    study.optimize(evaluate)?;
    expect_count(&study, TRIALS, "the resumed handle")?;

    // The numbering continued rather than restarting: trial `SEGMENT` is the
    // first one the second handle ran, and it exists exactly once.
    let view = study.view()?;
    let first_after = TrialNumber::new(SEGMENT);
    if view.get(first_after).is_none() {
        return Err(Error::Conflict(format!(
            "the resumed study has no trial #{SEGMENT}: the numbering restarted"
        )));
    }

    study
        .best_trial()?
        .ok_or_else(|| Error::NotFound("the resumed study has no best trial".to_owned()))
}
def run_in_two_halves(path: str) -> atune.FrozenTrial:
    """Runs the study in two halves, with the handle dropped in between.

    The second handle knows only the file. Everything else — the seed, the
    direction, the declared space, the seed protocol, the trial numbering, and any
    state the sampler or scheduler persisted — comes back out of storage, which is
    why a resumed study cannot silently disagree with the one it continues.
    """
    # First handle: create the study in the journal file and spend half the
    # budget. Dropping it at the end of this scope is the "interruption" — a
    # crashed worker, a pre-empted spot instance, a laptop closing.
    study = atune.create_study(
        direction="minimize", storage=path, name=STUDY, seed=SEED
    )
    study.optimize(objective, n_trials=SEGMENT)
    expect_count(study, SEGMENT, "the first handle")
    del study

    # Second handle: `load_study` on the same path. `n_trials` counts the trials
    # this call runs, so the study reaches TRIALS in total.
    study = atune.load_study(path, name=STUDY)
    study.optimize(objective, n_trials=TRIALS - SEGMENT)
    expect_count(study, TRIALS, "the resumed handle")

    # The numbering continued rather than restarting: trial `SEGMENT` is the first
    # one the second handle ran, and it exists exactly once.
    numbers = [trial.number for trial in study.trials]
    if numbers.count(SEGMENT) != 1:
        msg = (
            f"the resumed study has no single trial #{SEGMENT}: the numbering restarted"
        )
        raise AssertionError(msg)

    best = study.best_trial
    if best is None:
        msg = "the resumed study has no best trial"
        raise AssertionError(msg)
    return best

That is the whole of resuming from the caller's side — open the same storage, keep going. In Rust the trial budget counts trials in the study, so a study of 100 trials that stopped at 60 needs the same budget of 100, not 40; Python's optimize(n_trials=…) counts the trials that call runs, so the resumed handle asks for the 40 that are left. Python can also leave the count out entirely — optimize(objective, n_trials=None, timeout=1800) runs for half an hour and then stops, and Study.stop() from inside an objective or an optimize(callbacks=…) callback ends the run at the trial that decides it has seen enough. Both mean "start no new trial after this": the trial in flight always finishes and is recorded, so a resumed study never inherits a half-written one.

Reopening it, in either language

Rust reopens with open_storage(spec) + builder().load(id) — the listing passes the constant id 0 a single-study journal always has, and StudyLocator is the lookup for storages that hold more; Python with atune.load_study(spec), or create_study(..., load_if_exists=True) when the same script both creates and resumes. Core storage can contain multiple studies; these adapters enforce one study per spec through the atomic catalog. They never probe a guessed numeric ID, and create_study without load_if_exists still raises on a non-empty storage rather than quietly starting a second study.

Prove the resume was exact

A resume that silently loses state is worse than a crash, because you keep the number. So the example does not illustrate resuming — it checks it, by running the same study both ways and comparing:

/// Fails unless two trials are the same answer, to the bit.
///
/// Values and parameters are compared as `f64::to_bits`, not with `==`: two
/// numbers that print identically can still differ in the last place, and "the
/// resumed run agrees to six decimals" is not the claim being made.
fn require_identical(resumed: &FrozenTrial, control: &FrozenTrial) -> Result<()> {
    if resumed.number != control.number {
        return Err(Error::Conflict(format!(
            "the resumed study's best is #{}, the uninterrupted study's is #{}",
            resumed.number.get(),
            control.number.get()
        )));
    }
    let (left, right) = (value_of(resumed)?, value_of(control)?);
    if left.to_bits() != right.to_bits() {
        return Err(Error::Conflict(format!(
            "best values differ: resumed {left:?} (bits {:#018x}) vs uninterrupted {right:?} \
             (bits {:#018x})",
            left.to_bits(),
            right.to_bits()
        )));
    }
    for (name, value) in &resumed.params {
        let other = control.params.get(name).ok_or_else(|| {
            Error::Conflict(format!("the uninterrupted study's best has no `{name}`"))
        })?;
        if !same_bits(*value, *other) {
            return Err(Error::Conflict(format!(
                "parameter `{name}` differs: resumed {value} vs uninterrupted {other}"
            )));
        }
    }
    Ok(())
}

/// `true` when two parameter values are identical — floats by their bits.
fn same_bits(left: ParamValue, right: ParamValue) -> bool {
    match (left, right) {
        (ParamValue::F64(left), ParamValue::F64(right)) => left.to_bits() == right.to_bits(),
        (left, right) => left == right,
    }
}
def require_identical(resumed: atune.FrozenTrial, control: atune.FrozenTrial) -> None:
    """Raises unless two trials are the same answer, to the bit.

    Values and parameters are compared through ``float.hex()``, not with ``==``:
    two numbers that print identically can still differ in the last place, and
    "the resumed run agrees to six decimals" is not the claim being made. It is
    the same comparison the Rust arm makes with ``f64::to_bits``.
    """
    if resumed.number != control.number:
        msg = (
            f"the resumed study's best is #{resumed.number}, "
            f"the uninterrupted study's is #{control.number}"
        )
        raise AssertionError(msg)
    left, right = value_of(resumed), value_of(control)
    if left.hex() != right.hex():
        msg = f"best values differ: resumed {left!r} ({left.hex()}) vs uninterrupted {right!r} ({right.hex()})"
        raise AssertionError(msg)
    for name, value in resumed.params.items():
        if name not in control.params:
            msg = f"the uninterrupted study's best has no `{name}`"
            raise AssertionError(msg)
        other = control.params[name]
        if not same_bits(value, other):
            msg = (
                f"parameter `{name}` differs: resumed {value} vs uninterrupted {other}"
            )
            raise AssertionError(msg)


def same_bits(left: object, right: object) -> bool:
    """``True`` when two parameter values are identical — floats by their bits."""
    if isinstance(left, float) and isinstance(right, float):
        return left.hex() == right.hex()
    return left == right

Twenty trials, reopen, twenty more, against forty uninterrupted in memory. The comparison is on f64::to_bits, not ==: two numbers that print identically can still differ in the last place, and "agrees to six decimals" is not the claim.

This is worth stealing for your own storage backend. If you implement Storage, the conformance suite is the real test — but a resume-equals-uninterrupted check is a small sanity assertion on top of it.

Add workers

Three levels, in increasing order of how much can go wrong:

Threads in one process. parallelism(n) on the builder. Same storage, same process, n objectives evaluated at once. For a stateless sampler the study is still reproducible trial-for-trial: the worker count is not an input to which configurations get evaluated.

Processes on one machine. Point each at the same journal file. Trial numbers are assigned race-safely by storage, so two workers never claim the same number. The CLI does this natively — several atune run invocations against one --study path are one study.

Machines. One process runs atune serve --storage <spec>, and the others use an atune://host:port storage spec. The wire protocol mirrors the Storage trait operation for operation, and the proof it does is that the same conformance suite that tests the local backends passes over TCP unchanged.

atune serve is a plaintext TCP endpoint with no authentication and no TLS. It defaults to 127.0.0.1:7878 and warns on stderr before serving when --bind selects a non-loopback address. A reachable peer can read and modify every study, so a warning is not an access control. See the security policy for the full boundary and incident runbook.

For a remote deployment, establish the outer controls before starting the listener:

  1. Use a private VPN such as Tailscale or WireGuard, then restrict the host firewall and any network/cloud ACL to the expected worker addresses and port. Deny public ingress; a private LAN without those rules is not an authorization boundary.
  2. For client identity, bind atune to loopback and put a layer-4 mTLS proxy in front. The proxy must validate and authorize client certificates before forwarding the raw TCP stream. Atune still has no auth/TLS on its local hop.
  3. Set --max-connections to the number of workers the host can support. This bounds admitted handlers, but does not stop an authorized or reachable peer from changing study data and is not complete denial-of-service protection.

Keep Journal (*.atj) and SQLite (*.db, *.sqlite, *.sqlite3) files in a private directory owned by the service account, with restrictive modes (for example umask 077). Protect SQLite -wal/-shm files and backups the same way; use an encrypted volume and encrypted backup transfer when study data is confidential. The network wrapper does not encrypt storage at rest.

The Rust/Python objective and every CLI/Oniro subprocess are operator-selected code. They inherit the launcher's privileges, environment, filesystem, and network access; argv substitution has no shell, but it is not a sandbox. Put untrusted work in an external OS/container/VM/job boundary with a dedicated unprivileged account and explicit CPU, memory, process, disk, wall-time, and network-egress limits. Atune timeouts, output caps, and Oniro caps are local guards, not a complete resource or privilege boundary. Keep credentials out of command arguments, storage specs, trial records, and output; the service banner and bounded child stderr can appear in logs or study data.

Assign a deployment owner before launch. On suspected exposure, stop the server, deny the port, revoke VPN/mTLS access, rotate affected credentials, and preserve a read-only copy of evidence. Atune supplies no credential revocation or rollback workflow: restore a verified backup to a new path and validate it before cutover. See the CLI reference for the flags and the feature reference for the remote feature that provides it.

Structural costs to plan for

The Journal backend publishes a complete replacement image for every mutation: it copies the current canonical image into a generation-bound ticket, appends one record, and atomically renames the ticket over the Journal. If the current image is B bytes, one successful mutation copies O(B) bytes before its new record is added. Across N mutations, the copied volume is Σ O(B_i); when the image grows roughly linearly, that cumulative volume is O(N²). This is write amplification, not a throughput promise. The journal_bytes_staged and journal_canonical_size metrics expose the staged and published sizes for an operator's workload. A long local-disk run should prefer SQLite; Journal remains the shared-filesystem option.

Both Journal durability modes sync the complete staged image before its atomic rename. Durability::Record additionally requests a containing-directory sync after publication where the platform supports it; Durability::Os leaves that metadata to filesystem policy. These are local durability requests, not power-loss evidence for a particular filesystem or NFS export.

journal_lock_wait reports the append-lock acquisition call after the process-local mutex has been acquired, including filesystem retries, as one terminal duration event. It does not measure time queued on that in-process mutex. A successful acquisition after a holder is contended; a terminal timeout is failed, even though it waited. It is not a per-retry count.

An active lifecycle trial using Journal renews its owner lease every one second in the system build. This is separate from Journal's default heartbeat_interval_millis = 30_000: that setting supplies the default 60-second stale-trial fail-over grace (2 × interval), unless stale_after_millis overrides it. It is not the live-trial renewal cadence. Journal persists each lifecycle renewal; renewal coalescing is not enabled. With P live Journal trials, the configured cadence nominally adds P durable renewal mutations per second, each with the full-image staging shape above.

Finalization recovery has a separate rule: it may reclaim a held claim only after the exact claim remains unchanged for three seconds of local monotonic observation, with nominal 100 ms polling. This recovery grace does not use Journal's heartbeat setting or wall-clock timestamps. Direct finalizer admission remains nonblocking; a terminal call's pending-work drain can invoke the same recovery path. A live but stalled callback can lose its ownership fence after the grace; its later guarded writes are rejected. Claim-acquisition errors are not retried automatically. In builds without system, including browser builds, recovery leaves a held claim in place and fails closed.

Finalization callback effects may run more than once across a crash or claim takeover. Give external effects stable identities and make them idempotent; exactly-once external effects are not promised.

lease_renewal_latency and active_trial_leases cover those lifecycle trial leases only. The separate filesystem append-lock sidecar renewer is not represented by either metric.

Observer refreshes have their own full-history cost. StudyReader::refresh, atune report, and every atune top refresh read and hydrate the full StudySnapshot, including requested side state and pending finalizations; they do not incrementally update only the newest trial. The observer then clones and orders the trial list, scans history, and rebuilds the immediate-mode view. If a consumer requests the snapshot's Pareto front, multi-objective ranking retains its O(M·N²) comparison cost for N comparable trials and M objectives. atune report pays its full summary once; atune top pays the full summary on each refresh interval. A front view pays its comparison cost when requested. The raw StudyView remains the lower-overhead API for samplers and schedulers, but no observer refresh-throughput claim is made.

What breaks when you scale

  • A history-dependent sampler or scheduler stops being reproducible trial-for-trial. TPE, DEHB, NSGA-II and every pruner see a different history under a different interleaving. The parameters of trial n are still a pure function of the seed for a stateless sampler; with a stateful one, what changes is which trials the model was fitted to. Determinism is precise about the boundary.
  • A journal grows without bound. It is append-only by design — that is what makes it crash-safe and readable by another tool — and each mutation stages the full current image. A very long local-disk study is better served by the sqlite backend; use the Journal when shared-filesystem reach matters.
  • A lost worker costs its trials in flight, not the study. Those trials are left Running; nothing else is affected.
  • atune serve is a single point of failure while it is up. It is a convenience for sharing a study, not a replicated store.

Next