Storage¶
A study lives in storage, not in the handle you built it with. That is the single idea this page exists to install, because almost everything else follows from it: a study survives the process that created it, several workers can drive one study at once, and there is no coordinator process — storage is the only thing workers coordinate through.
What storage is responsible for¶
Storage is not a passive log. Five of its duties are correctness requirements that every backend must meet, and they are what make coordinator-free workers safe:
- Trial numbers. Storage assigns them: contiguous from 0, per study, race safe. Two workers asking for a trial at the same instant get two different numbers with no gap and no duplicate. Seeds are derived from the number, so a gap or a repeat would break reproducibility, not just tidiness.
- Study identity, and the configuration under it. The seed, the directions, the declared space, the resource unit and the metric names come back exactly as they went in — which is what lets a second process reproduce the first one's sampling.
- Exactly one owner per running trial. Study execution uses the atomic lifecycle capability: reserve, claim, and renew return an owner/epoch token, and every ownership-scoped write checks that token. Raw CRUD retains its compare-and-set interface for compatibility, but Study never emulates an atomic claim with separate raw writes.
- Terminal work is recoverable. Terminal outcome, error text, seam state, and retry/requeue intent are committed with a durable idempotent outbox. Replaying completed work is a no-op.
- Trial budgets are global. Reserving and creating a trial under
max_trialsis one atomic storage operation. The limit counts durable trial records across handles and processes, including a record whose later startup fails.
Storage never reads the operating-system clock implicitly. Backends use an
injected Clock for storage-owned timestamps, while caller timestamps remain
explicit. Tests can therefore drive time deterministically. Lease expiry never
compares wall-clock readings from two different workers.
Study names are deliberately not unique. Built-in backends expose the optional
StudyCatalog capability for atomic create-if-empty, deterministic listing,
and exact-name lookup; duplicate names return a typed conflict rather than an
arbitrary identity. Consumers such as the CLI and Python layer their
one-study-per-spec policy on that catalog. A third-party backend may omit the
capability, but adapters then fail explicitly instead of probing numeric IDs.
Storage also cannot promise durability beyond what the chosen backend provides.
One storage spec, one study¶
A storage spec names exactly one study. That is the rule the CLI and the Python binding hold you to, and it is worth stating on its own because a reader arriving from a tool whose storage is a registry of studies will otherwise assume the wrong model and find out by miscounting trials.
There is no registry because nothing needs one. A study's identity is the spec you
opened, so atune run --study s.atj needs no --study-name, an
atune://host:port worker needs no lookup step, and there is no name to collide
over. Four things follow, and each is a promise rather than a limitation:
- Creating a study in a non-empty storage is an error, not a second study. The empty check and the insertion are one atomic operation, so two processes racing to create cannot both win — and neither gets a study that renumbers its trials from zero while looking like a resume.
- Reopening is a separate verb:
atune.load_study(spec)from Python,create_study(..., load_if_exists=True)for a script that both creates and resumes,StudyLocatorplusbuilder().load(id)from Rust. - A name is a check, not a lookup. Passing one verifies that the study in the spec is the study you meant, and mismatching says which name is actually there.
- There is no
get_all_study_namesand nodelete_study, because there is nothing to enumerate and the file is the study. Removing a study is removing its storage.
What is true one layer down, since the two are easy to conflate: the Storage
trait itself is a multi-study CRUD seam and always was — a StudyId addresses a
study, and a backend may hold several. The one-study rule is an adapter policy
layered on the optional StudyCatalog capability described above, which is what
makes the empty check atomic and the name lookup exact. A backend that holds several
studies is legal, and the Python binding refuses it explicitly rather than choosing
one; a backend that omits the catalog fails with a typed unsupported error rather
than probing StudyId(0). Migrating from a tool with a study registry starts at
Migrate from Optuna.
The backends¶
Every backend answers the same questions the same way — the conformance suite further down this page is what holds them to it — so the choice between them is operational, never semantic. No backend gives you a different study; each gives you a different set of costs. The storage catalogue names each one's Rust type and the cargo feature that gates it, and not all of them are on in a default build. What it cannot tell you is which one you want; these are structural costs, not throughput guarantees:
| Backend | Reach | What it buys | What it costs |
|---|---|---|---|
| In memory | Threads in one process | Nothing to configure, nothing left behind, and it is the reference the others are calibrated against | The study dies with the process |
| Journal file | Processes on a shared filesystem | Resume is replay; no daemon; a study you can grep; crash tolerance by construction; designed for NFS with a mount-specific live test |
Reading replays the log; broad summaries still materialize history, and there are no indexed queries; an untested NFS export is not certified |
| SQLite | Many workers on one host | Indexed targeted queries and real SQL for tooling | Broad summaries still materialize required history; it does not work over NFS, and the bundled engine costs a C compile |
| Remote | Many machines, no shared filesystem | Wraps any other backend and serves it over TCP | A round trip per operation, and no authentication and no encryption — it binds to localhost unless you say otherwise |
The remote backend wrapping any other is worth a second look: remote-over-journal and remote-over-SQLite are both valid stacks, so "which backend" and "how do workers reach it" are independent choices.
Naming one¶
The Rust facade and CLI name each backend with a single string. Python's
create_study(storage=…) and load_study(…) accept the in-memory, journal and
SQLite forms below; the Python wheel does not enable the facade's remote
feature, so atune:// is not a Python locator:
| Spec | Backend |
|---|---|
:memory: |
In memory. This is the default when you say nothing. |
study.atj |
A journal file — the extension is what selects it |
study.db, study.sqlite, study.sqlite3 |
SQLite |
atune://host:port |
A remote server (Rust and CLI only) |
Naming a backend your build does not have is a clear error saying which cargo feature is missing, not a mysterious failure — and never a silent fallback.
Resuming a study¶
Because the study is in storage, stopping is not losing. A later handle opens the same file, loads the study by id, and continues: the trial numbering carries on where it left off, the seed and the space come back out of storage, and whatever state the sampler or scheduler persisted is restored.
The example below does exactly that, and then proves it. It runs one study in two halves across two handles on a journal file, runs the same study once uninterrupted in memory as a control, and requires the two to agree on the best trial's number, its value and every parameter, to the bit. "Resume works" is easy to demonstrate and easy to get subtly wrong — a resumed study that quietly restarts its seed still looks right — so the example asserts rather than prints:
/// Runs the study in two halves, with the handle dropped in between.
///
/// The second handle knows only the file and the study id. Everything else — the
/// seed, the direction, the declared space, the seed protocol, the trial
/// numbering, and any state the sampler or scheduler persisted — comes back out
/// of storage, which is why a resumed study cannot silently disagree with the one
/// it continues.
fn run_in_two_halves(path: &Path) -> Result<FrozenTrial> {
let spec = path.to_string_lossy().into_owned();
// First handle: create the study in the journal file and spend half the
// budget. Dropping it at the end of this scope is the "interruption" — a
// crashed worker, a pre-empted spot instance, a laptop closing.
{
let study = Study::builder()
.storage(atune::open_storage(&spec)?)
.budget(Budget::trials(SEGMENT))
.create(StudyConfig::new(STUDY).with_seed(SEED))?;
study.optimize(evaluate)?;
expect_count(&study, SEGMENT, "the first handle")?;
}
// Second handle: a fresh `open_storage` on the same path, then `load` rather
// than `create`. The budget counts the trials already in the study, so
// `trials(TRIALS)` means "run up to TRIALS", not "run TRIALS more".
let study = Study::builder()
.storage(atune::open_storage(&spec)?)
.budget(Budget::trials(TRIALS))
.load(ONLY_STUDY)?;
study.optimize(evaluate)?;
expect_count(&study, TRIALS, "the resumed handle")?;
// The numbering continued rather than restarting: trial `SEGMENT` is the
// first one the second handle ran, and it exists exactly once.
let view = study.view()?;
let first_after = TrialNumber::new(SEGMENT);
if view.get(first_after).is_none() {
return Err(Error::Conflict(format!(
"the resumed study has no trial #{SEGMENT}: the numbering restarted"
)));
}
study
.best_trial()?
.ok_or_else(|| Error::NotFound("the resumed study has no best trial".to_owned()))
}
def run_in_two_halves(path: str) -> atune.FrozenTrial:
"""Runs the study in two halves, with the handle dropped in between.
The second handle knows only the file. Everything else — the seed, the
direction, the declared space, the seed protocol, the trial numbering, and any
state the sampler or scheduler persisted — comes back out of storage, which is
why a resumed study cannot silently disagree with the one it continues.
"""
# First handle: create the study in the journal file and spend half the
# budget. Dropping it at the end of this scope is the "interruption" — a
# crashed worker, a pre-empted spot instance, a laptop closing.
study = atune.create_study(
direction="minimize", storage=path, name=STUDY, seed=SEED
)
study.optimize(objective, n_trials=SEGMENT)
expect_count(study, SEGMENT, "the first handle")
del study
# Second handle: `load_study` on the same path. `n_trials` counts the trials
# this call runs, so the study reaches TRIALS in total.
study = atune.load_study(path, name=STUDY)
study.optimize(objective, n_trials=TRIALS - SEGMENT)
expect_count(study, TRIALS, "the resumed handle")
# The numbering continued rather than restarting: trial `SEGMENT` is the first
# one the second handle ran, and it exists exactly once.
numbers = [trial.number for trial in study.trials]
if numbers.count(SEGMENT) != 1:
msg = (
f"the resumed study has no single trial #{SEGMENT}: the numbering restarted"
)
raise AssertionError(msg)
best = study.best_trial
if best is None:
msg = "the resumed study has no best trial"
raise AssertionError(msg)
return best
Both tabs do the same thing, and neither supplies an id it had to look up: the
Rust arm passes the constant id 0 that a single-study journal always has, the
Python arm passes none at all — the spec is the identity, exactly as
above.
Several workers, one study¶
In one process, threads share a handle and are serialised inside it.
Across processes on a shared filesystem, the journal backend serialises
writers with a renewable fenced directory lease, not flock, because POSIX
record locks are unreliable over NFS. Readers never take the lease. A claim
contains a nonce and renewable epoch; a contender may reclaim it only after the
same complete claim stays unchanged for one grace period on the contender's
local monotonic clock. A generation-bound append ticket prevents a late old
writer from publishing after takeover. Pre-lease lock files are diagnosed but
never auto-broken; coordinated cleanup is required before mixing old and new
writers. The real-NFS ignored test is the evidence boundary for a mount's
exclusive-create and atomic-rename behavior.
Across machines without a shared filesystem, the remote backend is the
answer, and SQLite is not: its write-ahead log is a shared-memory protocol and
does not work over NFS. That is why both file backends exist. The server side of
it is atune serve, which takes a spec of its
own — so what the workers talk to and what actually holds the study stay two
separate decisions.
Remote telemetry separates client-side serialization from service time. The
atune::metrics target emits remote_queue_time as whole microseconds waiting
for that handle's socket mutex and remote_request_latency as whole
microseconds after acquisition, including transport, backend work, and any
lifecycle reconnect/replay. Their shared operation value joins the two events;
it is not a metric label. Lifecycle traffic uses a separate lazy control socket,
so a blocked data-page read cannot hold its channel, although calls within each
channel still serialize. remote_page_restart counts the one bounded retry
after an immutable page session expires or is evicted. remote_replay_eviction
counts completed lifecycle results removed by the entry or byte cap: eviction
shortens replay availability and never authorizes automatic re-execution. These
events expose a workload's shape; atune claims no remote latency or throughput
threshold without controlled measurement.
Whichever you pick, running trials write heartbeats. Failover requires the same owner/epoch/heartbeat observation to remain unchanged for one sweeper-local grace period, then atomically records one failed outcome and durable replacement intent when requeue is enabled. A renewal racing that commit leaves either the live owner intact or exactly one complete stale failure.
The Journal's write shape is intentionally visible in its cost model. Every
mutation stages the full current JSONL image and atomically replaces the
canonical path. For a current image of B bytes, that mutation copies
O(B) bytes; over N mutations the cumulative copied volume is
Σ O(B_i), potentially O(N²) as the image grows. journal_bytes_staged and
journal_canonical_size report those sizes; they do not imply a throughput
bound. Prefer SQLite for a long local-disk run, and retain Journal for the
shared-filesystem case.
Both Journal durability modes sync the complete staged image before its
atomic rename. Durability::Record also requests a sync of the containing
directory after publication where the platform supports it; Durability::Os
leaves that directory metadata to filesystem policy. These calls describe the
local requests, not proof of power-loss behavior for a particular filesystem or
NFS export.
journal_lock_wait reports the append-lock acquisition call after the
process-local mutex has been acquired, including filesystem retries, as one
terminal duration event. It does not measure time queued on that in-process
mutex. A successful acquisition after a holder is contended; a terminal
timeout is failed, even though it waited. It is not a per-retry count.
The live-trial lease and stale-trial policy use different cadences. A system
build renews each active Journal lifecycle lease every one second. Journal's default
heartbeat_interval_millis is 30 seconds and only derives the 60-second
stale-trial grace (twice that interval), unless stale_after_millis is set.
The 30-second policy value does not slow the one-second durable renewals, and
Journal persists each renewal without coalescing. With P live Journal trials,
the configured cadence nominally adds P durable renewal mutations per second,
each with the full-image staging shape above.
lease_renewal_latency and active_trial_leases describe lifecycle trial
leases. They do not instrument the journal append-lock sidecar renewer, whose
filesystem lease is a separate mechanism.
Read-only observers also pay a structural refresh cost. StudyReader::refresh
and the atune top loop rebuild a full hydrated StudySnapshot, including
trial history, requested side state, and pending finalizations, on each
refresh. The observer then clones/orders trials, scans history, rebuilds
immediate-mode geometry. If a consumer requests the snapshot's Pareto front,
the O(M·N²) multi-objective comparison remains for N comparable trials and
M objectives. atune report pays its full summary once; atune top pays the
full summary on every refresh. A front view pays its comparison cost when
requested. These are structural costs, not benchmarked throughput claims.
Policy-aware growth replay consumes the policy map carried by that same
snapshot. Exact snapshot backends enumerate every persisted policy slot in the
read; a source-compatible BestEffort fallback cannot enumerate named state,
so it can replay only names already known from the configured space or trial
records and may miss a policy-only slot written concurrently. That is a
completeness boundary, not a throughput claim.
Hydration's batching claim has a deliberate boundary. Exact built-in backends
(Memory, Journal, and SQLite) and remote over an exact built-in backend fulfill
one read_snapshot call from one backend read epoch (remote may page that
immutable result). A custom backend that relies on the source-compatible
default is marked BestEffort:
each pass makes one sync trial-list read, one get_state call for every
concrete requested scope, one list_outbox call when its lifecycle capability
is available, and one get_trial call for every listed outbox record. If a
returned trial introduces a policy name absent from the immutable
configuration, hydration makes at most one second pass with the discovered
names. This is a bounded call shape, not a throughput guarantee; the marker
still permits concurrent writes between those separate reads. StudyReader::user_attrs()
is a separate study_user_attrs read, outside the snapshot epoch, so combining
it with a snapshot can observe mixed-epoch metadata.
Resume and scale is the recipe.
Writing your own backend¶
The storage seam is a trait with twenty methods and a written contract, and the contract is not left as prose: a conformance suite ships behind a cargo feature, and running it against your backend produces one test per invariant, so a failure names the invariant you broke rather than saying that the suite failed. It covers numbering under real thread contention, the transition matrix, immutability, the parameter-write gate, report ordering, state blobs and the synchronisation cursor.
All four backends in this repository are themselves customers of it, and one of them runs the whole suite over TCP. Extending: storage walks through it.
Where to go next¶
| If you want to | Go to |
|---|---|
| Keep a study across runs and machines | Resume and scale |
| Read a stored study | Inspect a study |
| Look a backend up by type, with its feature gate | Storage catalogue |
| Know what is written into a journal file, record by record | Journal format |
| Open a study in Optuna's dashboard | Optuna interop |
| Implement the trait yourself | Extending: storage |
| Understand what resuming preserves, and why | Determinism |