Resume and scale a study¶
Goal: survive an interrupted run, and add workers — threads, processes or machines — without a coordinator.
Both halves of this page are the same idea: a study lives in storage, not in the handle that created it. Point a study at a file and it outlives the process. Point several workers at the same storage and they are the same study. There is no scheduler daemon and no server to stand up — storage is the only thing workers coordinate through.
Give the study a file¶
An in-memory study dies with the process. A journal-backed one does not: it is an append-only file that records every operation, so a later handle re-reads it, continues the trial numbering, restores whatever state the sampler and scheduler persisted, and carries on.
/// Runs the study in two halves, with the handle dropped in between.
///
/// The second handle knows only the file and the study id. Everything else — the
/// seed, the direction, the declared space, the seed protocol, the trial
/// numbering, and any state the sampler or scheduler persisted — comes back out
/// of storage, which is why a resumed study cannot silently disagree with the one
/// it continues.
fn run_in_two_halves(path: &Path) -> Result<FrozenTrial> {
let spec = path.to_string_lossy().into_owned();
// First handle: create the study in the journal file and spend half the
// budget. Dropping it at the end of this scope is the "interruption" — a
// crashed worker, a pre-empted spot instance, a laptop closing.
{
let study = Study::builder()
.storage(atune::open_storage(&spec)?)
.budget(Budget::trials(SEGMENT))
.create(StudyConfig::new(STUDY).with_seed(SEED))?;
study.optimize(evaluate)?;
expect_count(&study, SEGMENT, "the first handle")?;
}
// Second handle: a fresh `open_storage` on the same path, then `load` rather
// than `create`. The budget counts the trials already in the study, so
// `trials(TRIALS)` means "run up to TRIALS", not "run TRIALS more".
let study = Study::builder()
.storage(atune::open_storage(&spec)?)
.budget(Budget::trials(TRIALS))
.load(ONLY_STUDY)?;
study.optimize(evaluate)?;
expect_count(&study, TRIALS, "the resumed handle")?;
// The numbering continued rather than restarting: trial `SEGMENT` is the
// first one the second handle ran, and it exists exactly once.
let view = study.view()?;
let first_after = TrialNumber::new(SEGMENT);
if view.get(first_after).is_none() {
return Err(Error::Conflict(format!(
"the resumed study has no trial #{SEGMENT}: the numbering restarted"
)));
}
study
.best_trial()?
.ok_or_else(|| Error::NotFound("the resumed study has no best trial".to_owned()))
}
def run_in_two_halves(path: str) -> atune.FrozenTrial:
"""Runs the study in two halves, with the handle dropped in between.
The second handle knows only the file. Everything else — the seed, the
direction, the declared space, the seed protocol, the trial numbering, and any
state the sampler or scheduler persisted — comes back out of storage, which is
why a resumed study cannot silently disagree with the one it continues.
"""
# First handle: create the study in the journal file and spend half the
# budget. Dropping it at the end of this scope is the "interruption" — a
# crashed worker, a pre-empted spot instance, a laptop closing.
study = atune.create_study(
direction="minimize", storage=path, name=STUDY, seed=SEED
)
study.optimize(objective, n_trials=SEGMENT)
expect_count(study, SEGMENT, "the first handle")
del study
# Second handle: `load_study` on the same path. `n_trials` counts the trials
# this call runs, so the study reaches TRIALS in total.
study = atune.load_study(path, name=STUDY)
study.optimize(objective, n_trials=TRIALS - SEGMENT)
expect_count(study, TRIALS, "the resumed handle")
# The numbering continued rather than restarting: trial `SEGMENT` is the first
# one the second handle ran, and it exists exactly once.
numbers = [trial.number for trial in study.trials]
if numbers.count(SEGMENT) != 1:
msg = (
f"the resumed study has no single trial #{SEGMENT}: the numbering restarted"
)
raise AssertionError(msg)
best = study.best_trial
if best is None:
msg = "the resumed study has no best trial"
raise AssertionError(msg)
return best
That is the whole of resuming from the caller's side — open the same storage, keep
going. In Rust the trial budget counts trials in the study, so a study of 100
trials that stopped at 60 needs the same budget of 100, not 40; Python's
optimize(n_trials=…) counts the trials that call runs, so the resumed handle
asks for the 40 that are left. Python can also leave the count out entirely —
optimize(objective, n_trials=None, timeout=1800) runs for half an hour and then
stops, and Study.stop() from inside an objective or an
optimize(callbacks=…) callback ends the run at the trial that decides it has
seen enough. Both mean "start no new trial after this": the trial in flight
always finishes and is recorded, so a resumed study never inherits a half-written
one.
Reopening it, in either language
Rust reopens with open_storage(spec) + builder().load(id) — the listing
passes the constant id 0 a single-study journal always has, and
StudyLocator is the lookup for storages that hold more; Python with
atune.load_study(spec), or create_study(..., load_if_exists=True) when the
same script both creates and resumes. Core storage can contain multiple studies;
these adapters enforce one study per spec through the atomic catalog. They never
probe a guessed numeric ID, and create_study without load_if_exists still
raises on a non-empty storage rather than quietly starting a second study.
Prove the resume was exact¶
A resume that silently loses state is worse than a crash, because you keep the number. So the example does not illustrate resuming — it checks it, by running the same study both ways and comparing:
/// Fails unless two trials are the same answer, to the bit.
///
/// Values and parameters are compared as `f64::to_bits`, not with `==`: two
/// numbers that print identically can still differ in the last place, and "the
/// resumed run agrees to six decimals" is not the claim being made.
fn require_identical(resumed: &FrozenTrial, control: &FrozenTrial) -> Result<()> {
if resumed.number != control.number {
return Err(Error::Conflict(format!(
"the resumed study's best is #{}, the uninterrupted study's is #{}",
resumed.number.get(),
control.number.get()
)));
}
let (left, right) = (value_of(resumed)?, value_of(control)?);
if left.to_bits() != right.to_bits() {
return Err(Error::Conflict(format!(
"best values differ: resumed {left:?} (bits {:#018x}) vs uninterrupted {right:?} \
(bits {:#018x})",
left.to_bits(),
right.to_bits()
)));
}
for (name, value) in &resumed.params {
let other = control.params.get(name).ok_or_else(|| {
Error::Conflict(format!("the uninterrupted study's best has no `{name}`"))
})?;
if !same_bits(*value, *other) {
return Err(Error::Conflict(format!(
"parameter `{name}` differs: resumed {value} vs uninterrupted {other}"
)));
}
}
Ok(())
}
/// `true` when two parameter values are identical — floats by their bits.
fn same_bits(left: ParamValue, right: ParamValue) -> bool {
match (left, right) {
(ParamValue::F64(left), ParamValue::F64(right)) => left.to_bits() == right.to_bits(),
(left, right) => left == right,
}
}
def require_identical(resumed: atune.FrozenTrial, control: atune.FrozenTrial) -> None:
"""Raises unless two trials are the same answer, to the bit.
Values and parameters are compared through ``float.hex()``, not with ``==``:
two numbers that print identically can still differ in the last place, and
"the resumed run agrees to six decimals" is not the claim being made. It is
the same comparison the Rust arm makes with ``f64::to_bits``.
"""
if resumed.number != control.number:
msg = (
f"the resumed study's best is #{resumed.number}, "
f"the uninterrupted study's is #{control.number}"
)
raise AssertionError(msg)
left, right = value_of(resumed), value_of(control)
if left.hex() != right.hex():
msg = f"best values differ: resumed {left!r} ({left.hex()}) vs uninterrupted {right!r} ({right.hex()})"
raise AssertionError(msg)
for name, value in resumed.params.items():
if name not in control.params:
msg = f"the uninterrupted study's best has no `{name}`"
raise AssertionError(msg)
other = control.params[name]
if not same_bits(value, other):
msg = (
f"parameter `{name}` differs: resumed {value} vs uninterrupted {other}"
)
raise AssertionError(msg)
def same_bits(left: object, right: object) -> bool:
"""``True`` when two parameter values are identical — floats by their bits."""
if isinstance(left, float) and isinstance(right, float):
return left.hex() == right.hex()
return left == right
Twenty trials, reopen, twenty more, against forty uninterrupted in memory. The
comparison is on f64::to_bits, not ==: two numbers that print identically can
still differ in the last place, and "agrees to six decimals" is not the claim.
This is worth stealing for your own storage backend. If you implement Storage,
the conformance suite is the real test — but a
resume-equals-uninterrupted check is a small sanity assertion on top of it.
Add workers¶
Three levels, in increasing order of how much can go wrong:
Threads in one process. parallelism(n) on the builder. Same storage, same
process, n objectives evaluated at once. For a stateless sampler the study is
still reproducible trial-for-trial: the worker count is not an input to which
configurations get evaluated.
Processes on one machine. Point each at the same journal file. Trial numbers
are assigned race-safely by storage, so two workers never claim the same number.
The CLI does this natively — several atune run invocations against one --study
path are one study.
Machines. One process runs atune serve --storage <spec>, and the others use
an atune://host:port storage spec. The wire protocol mirrors the Storage trait
operation for operation, and the proof it does is that the same conformance suite
that tests the local backends passes over TCP unchanged.
atune serve is a plaintext TCP endpoint with no authentication and no TLS.
It defaults to 127.0.0.1:7878 and warns on stderr before serving when
--bind selects a non-loopback address. A reachable peer can read and modify
every study, so a warning is not an access control. See the
security policy
for the full boundary and incident runbook.
For a remote deployment, establish the outer controls before starting the listener:
- Use a private VPN such as Tailscale or WireGuard, then restrict the host firewall and any network/cloud ACL to the expected worker addresses and port. Deny public ingress; a private LAN without those rules is not an authorization boundary.
- For client identity, bind atune to loopback and put a layer-4 mTLS proxy in front. The proxy must validate and authorize client certificates before forwarding the raw TCP stream. Atune still has no auth/TLS on its local hop.
- Set
--max-connectionsto the number of workers the host can support. This bounds admitted handlers, but does not stop an authorized or reachable peer from changing study data and is not complete denial-of-service protection.
Keep Journal (*.atj) and SQLite (*.db, *.sqlite, *.sqlite3) files in a
private directory owned by the service account, with restrictive modes (for
example umask 077). Protect SQLite -wal/-shm files and backups the same
way; use an encrypted volume and encrypted backup transfer when study data is
confidential. The network wrapper does not encrypt storage at rest.
The Rust/Python objective and every CLI/Oniro subprocess are operator-selected code. They inherit the launcher's privileges, environment, filesystem, and network access; argv substitution has no shell, but it is not a sandbox. Put untrusted work in an external OS/container/VM/job boundary with a dedicated unprivileged account and explicit CPU, memory, process, disk, wall-time, and network-egress limits. Atune timeouts, output caps, and Oniro caps are local guards, not a complete resource or privilege boundary. Keep credentials out of command arguments, storage specs, trial records, and output; the service banner and bounded child stderr can appear in logs or study data.
Assign a deployment owner before launch. On suspected exposure, stop the
server, deny the port, revoke VPN/mTLS access, rotate affected credentials,
and preserve a read-only copy of evidence. Atune supplies no credential
revocation or rollback workflow: restore a verified backup to a new path and
validate it before cutover. See the CLI reference for
the flags and the feature reference for the
remote feature that provides it.
Structural costs to plan for¶
The Journal backend publishes a complete replacement image for every mutation:
it copies the current canonical image into a generation-bound ticket, appends
one record, and atomically renames the ticket over the Journal. If the current
image is B bytes, one successful mutation copies O(B) bytes before its new
record is added. Across N mutations, the copied volume is
Σ O(B_i); when the image grows roughly linearly, that cumulative volume is
O(N²). This is write amplification, not a throughput promise. The
journal_bytes_staged and journal_canonical_size metrics expose the staged
and published sizes for an operator's workload. A long local-disk run should
prefer SQLite; Journal remains the shared-filesystem option.
Both Journal durability modes sync the complete staged image before its
atomic rename. Durability::Record additionally requests a containing-directory
sync after publication where the platform supports it; Durability::Os leaves
that metadata to filesystem policy. These are local durability requests, not
power-loss evidence for a particular filesystem or NFS export.
journal_lock_wait reports the append-lock acquisition call after the
process-local mutex has been acquired, including filesystem retries, as one
terminal duration event. It does not measure time queued on that in-process
mutex. A successful acquisition after a holder is contended; a terminal
timeout is failed, even though it waited. It is not a per-retry count.
An active lifecycle trial using Journal renews its owner lease every one second
in the system build. This is separate from Journal's default
heartbeat_interval_millis = 30_000: that setting supplies the default
60-second stale-trial fail-over grace (2 × interval), unless
stale_after_millis overrides it. It is not the live-trial renewal cadence.
Journal persists each lifecycle renewal; renewal coalescing is not enabled. With
P live Journal trials, the configured cadence nominally adds P durable
renewal mutations per second, each with the full-image staging shape above.
Finalization recovery has a separate rule: it may reclaim a held claim only
after the exact claim remains unchanged for three seconds of local monotonic
observation, with nominal 100 ms polling. This recovery grace does not use
Journal's heartbeat setting or wall-clock timestamps. Direct finalizer
admission remains nonblocking; a terminal call's pending-work drain can invoke
the same recovery path. A live but stalled callback can lose its ownership
fence after the grace; its later guarded writes are rejected. Claim-acquisition errors are
not retried automatically. In builds without system, including browser
builds, recovery leaves a held claim in place and fails closed.
Finalization callback effects may run more than once across a crash or claim takeover. Give external effects stable identities and make them idempotent; exactly-once external effects are not promised.
lease_renewal_latency and active_trial_leases cover those lifecycle trial
leases only. The separate filesystem append-lock sidecar renewer is not
represented by either metric.
Observer refreshes have their own full-history cost. StudyReader::refresh,
atune report, and every atune top refresh read and hydrate the full
StudySnapshot, including requested side state and pending finalizations; they
do not incrementally update only the newest trial. The observer then clones and
orders the trial list, scans history, and rebuilds the immediate-mode view.
If a consumer requests the snapshot's Pareto front, multi-objective ranking
retains its O(M·N²) comparison cost for N comparable trials and M
objectives. atune report pays its full summary once; atune top pays the
full summary on each refresh interval. A front view pays its comparison cost
when requested. The raw StudyView remains the lower-overhead API for samplers
and schedulers, but no observer refresh-throughput claim is made.
What breaks when you scale¶
- A history-dependent sampler or scheduler stops being reproducible trial-for-trial. TPE, DEHB, NSGA-II and every pruner see a different history under a different interleaving. The parameters of trial n are still a pure function of the seed for a stateless sampler; with a stateful one, what changes is which trials the model was fitted to. Determinism is precise about the boundary.
- A journal grows without bound. It is append-only by design — that is what
makes it crash-safe and readable by another tool — and each mutation stages
the full current image. A very long local-disk study is better served by the
sqlitebackend; use the Journal when shared-filesystem reach matters. - A lost worker costs its trials in flight, not the study. Those trials are
left
Running; nothing else is affected. atune serveis a single point of failure while it is up. It is a convenience for sharing a study, not a replicated store.
Next¶
- Storage — the four duties every backend must meet.
- Inspect a study — reading a study back while it runs.
- Tune any program — the CLI path, which resumes by default.