Write a storage backend¶
Storage is the only thing workers coordinate through — there is no coordinator process — so its invariants are the ones a bug in your backend would corrupt a study through. They are subtle, mostly concurrent, and easy to almost get right.
Which is why the strongest thing this page has to tell you is not a walkthrough: the same conformance suite that validates the in-memory, journal, SQLite and TCP backends will validate yours, unchanged. The contract is not left as prose for you to interpret. It is a battery of executable checks, written once against a factory, and you can run it before you have written a single test of your own.
What you are signing up to¶
The base CRUD/read-model methods and a contract with four load-bearing promises: trial numbers are contiguous and race-safe; a study's configuration comes back exactly as it went in; every state change has exactly one winner; a finished trial is immutable. Reads are incremental through a cursor, and side-channel state is a typed, versioned blob rather than a junk drawer. Atomic lifecycle and study-catalog behavior are opt-in capabilities, so a backend can implement the base seam without pretending it can safely provide either.
The base Storage trait alone is not enough to run the normal study loop.
It supports direct CRUD and observation, but every Study-driven trial requires
the atomic TrialLifecycleStorage capability returned by
Storage::trial_lifecycle. Without it, ask/optimize fails with
Error::Unsupported; never reconstruct reserve, claim, renew, or fenced writes
from separate CRUD calls. A third-party backend intended for optimization must
implement that capability as well as the base trait.
Storage explains all of that from the user's side, and it is worth reading first — this page assumes it. One property is worth repeating here because it is a requirement on your code rather than a feature of the system: storage never reads a clock. Timestamps are handed in by the caller, which is what lets a test drive a time-bounded study deterministically.
Snapshot compatibility¶
Implementing the existing storage methods without overriding
Storage::read_snapshot remains source-compatible. The default result is
SnapshotConsistency::BestEffort, not an atomic observation. Per pass it does
one sync trial-list read, one get_state call per concrete requested scope,
and, when outbox reads are requested and the lifecycle capability is present,
bounded list_outbox_page reads followed by one get_trial call per listed
record. If paging is unsupported or a record exceeds the page budget, this
best-effort full-snapshot fallback uses list_outbox; its assembled result can
therefore be unbounded. Named-prefix state cannot be enumerated by this fallback.
StudySnapshot may make one
additional pass when returned trial parameters reveal policy names not present
in the immutable configuration; it does not retry discovery indefinitely.
Exact built-in backends (Memory, Journal, and SQLite) provide the one-epoch
Atomic result. A remote wrapper can retain that guarantee only when it wraps
an exact backend; a custom backend using the fallback remains BestEffort.
StudyReader::user_attrs() is a separate metadata read and is not covered by
the snapshot epoch.
Backends used for writable study load or finalization recovery must implement
the finalization claim operations (observe, acquire, renew and release) and the
guarded lifecycle operations. Recovery also requires bounded outbox paging.
Unsupported paging and records too large for the requested page fail closed;
they do not fall back to the unbounded list_outbox method. The full-snapshot
fallback above is a separate, best-effort observation path.
Finalization effects are at-least-once across crashes or claim takeover. Make external effects idempotent using stable effect identities. A live but stalled owner can lose its fence after recovery observes an unchanged claim for the recovery grace, so later guarded writes from that owner are rejected.
The suite is the guide¶
The suite's own module documentation is the step-by-step walkthrough, its code
examples are compiled and run by the repository's gate, and it lives at
crates/atune_core/src/storage/conformance.rs.
Rather than restate it, here is what it costs you and what it gives back.
It costs one dev-dependency and one feature. Depend on atune_core with the
conformance feature — as a dev-dependency, so the battery never ships in your
release builds:
atune_core = { version = "0.1", features = ["conformance"] } under
[dev-dependencies] — a path or git dependency until the first release, as
Install explains.
Name atune_core even if your crate otherwise depends on the facade. The facade's
conformance feature re-exports the module as atune::storage::conformance, so
run_all and the individual checks are reachable through it — but the macro below
is exported at atune_core's root and atune::storage_conformance_tests! does not
resolve. The facade's own SQLite suite names atune_core for that reason.
It asks for a factory. A closure returning a fresh, empty backend as a
Box<dyn Storage>. It is called once per check, so no check can be polluted by
another's state.
It gives you one test per invariant.
atune_core::storage_conformance_tests!(factory) expands to thirty-four #[test]
functions, each named after the invariant it pins. Put the call in its own module — the macro defines items with fixed names.
A single-test variant, run_all, exists for a smoke check.
That "one test per invariant" is the point of the shape. A failure reports the check, the invariant it broke, and what was observed instead, so you are told which promise your backend broke rather than that the storage suite failed.
Four backends already run it¶
| Backend | How it runs the suite |
|---|---|
| In memory | The reference implementation, and the oracle the suite is calibrated against |
| Journal file | The same battery over a real file, with its own extra tests for what the trait cannot express |
| SQLite | The suite unchanged, through the exported macro, exactly the way an out-of-tree backend would run it — that was the acceptance criterion for the slice that added it |
| Remote | The suite unchanged, over a real loopback TCP connection, against a fresh embedded server per check |
The remote row is the strongest evidence that the suite is genuinely backend-agnostic: every invariant is asserted across a socket, including the typed error cases and the non-finite float round trips, which are precisely what a subtly lossy translation layer would break.
The concurrent checks, and how to know they ran¶
Four checks spawn threads — atomic empty-catalog creation when that capability exists, race-free trial numbering, exactly-one-winner transitions, and interleaved writers under a sync cursor. They use barriers rather than sleeps, so their ordering is deterministic, and they exist unconditionally so that the macro's expansion never depends on your crate's feature set.
Their bodies compile only where there are threads to race with. On wasm32, or in
a --no-default-features build, they become no-ops. runs_concurrent_checks()
reports which build you are in, and the honest use of it is an assertion in
whichever build you consider your release gate. Thirty-four tests still pass in a
build where the concurrent checks assert nothing, and that is worse than a failure,
because it looks like success.
What the suite deliberately does not check¶
Durability, cross-process visibility and crash recovery. None of the three is
expressible through the trait, and each is backend-specific — the journal backend
brings its own tests for exactly this reason, and yours should too. The suite
also does not require unique study names. If the backend opts into
StudyCatalog, it must return a typed conflict for duplicates and atomically
create into an empty catalog; consumers such as the CLI can then enforce their
own one-study-per-file convention without guessing an ID.
How a user reaches your backend¶
By handing it over directly: the study builder takes an Arc<dyn Storage>, and
yours goes exactly where a built-in would.
What will not pick it up is a spec string. :memory:, study.atj,
study.db/study.sqlite and atune://host:port are matched by a closed
function in the facade. The CLI's --study flag goes through that function;
Python's create_study(storage=…) accepts the in-memory, journal and SQLite
forms, but the Python wheel does not enable the facade's remote feature and
therefore does not accept atune://host:port. So an out-of-tree backend is
usable from Rust and is not reachable from the CLI or from Python — worth
knowing before you plan a deployment around it.
Where to go next¶
| If you want to | Go to |
|---|---|
| Understand what storage is responsible for | Storage |
| See what the built-in backends buy and cost | Storage, and Resume and scale |
| Know what is written into a journal file | Journal format |
| Know which feature carries which backend | Feature reference |
| Ship it as a crate someone can depend on | Publish a plugin crate |
| Read the trait, method by method | Rust API |