Skip to content

Studies and trials

Two objects carry almost everything atune does. A study is one optimisation. A trial is one evaluation of one configuration inside it. Every other idea on this site — a sampler, a pruner, a storage backend, a seed — attaches to one of the two.

This page is language-neutral: it names the concept first and then how each language spells it.

A study

A study is a question ("which configuration minimises this?") plus everything needed to keep asking it: what to optimise and in which direction, where the answers are written down, who proposes the next configuration, who decides when a running trial has seen enough, and when to stop.

Concretely, a study handle binds all of this:

It owns Which means Default if you say nothing
A configuration Name, seed, one or more directions, an optional declared space, optional metric names, the unit its resources are counted in, and the multi-seed protocol Named by you; seed 0; a single Minimize; no declared space; unnamed metrics; resources counted in steps; one seed per trial
Storage Where trials are written and read back from In memory — the study dies with the process
A sampler Proposes the parameters of the next trial Random search
A scheduler Watches a running trial's intermediate reports and may stop, pause or fork it A scheduler that never intervenes
A clock The time source budgets and timestamps are read from The system clock
A budget When to stop creating trials Unbounded
A degree of parallelism How many trials are evaluated at once One

Here is a study, built in each language and then run:

let study = Study::builder()
    .parallelism(THREADS)
    .budget(Budget::trials(TRIALS))
    .create(StudyConfig::new(STUDY).with_seed(SEED))?;

study.optimize(|ctx| {
    let x = ctx.suggest_f64("x", -5.12..=5.12, Scale::Linear)?;
    let y = ctx.suggest_f64("y", -5.12..=5.12, Scale::Linear)?;
    Ok(rastrigin(&[x, y]).into())
})?;
study = atune.create_study(
    direction="minimize",
    seed=SEED,
    name=STUDY,
)
study.optimize(objective, n_trials=TRIALS, n_jobs=JOBS)

Two consequences of that list are worth stating plainly.

A study is not a handle. A study lives in storage; a handle is one worker's binding to it. Several handles — several threads, several processes, several machines — may drive the same study at once, and they coordinate only through storage. Drop every handle and the study is still there, provided storage outlives the process. That is what makes resuming and scaling possible at all.

The direction is a list, not a flag. A single-objective study is the case where that list has one element. Optimising two things at once is not a mode you switch on; it is the general case, and the single-objective API is the specialisation. Multi-objective and constraints picks that up.

A trial

A trial is one evaluation: a set of parameters, the value or values they produced, and the bookkeeping around them. Every trial in a study carries two identifiers, and confusing them is the most common way to get a wrong answer out of atune:

Identifier Scope Guarantees Use it for
Trial number Per study Contiguous from 0, assigned race-safely by storage, never reused Ordering, and seeding — this is the determinism key
Trial id Storage-global Opaque; may have gaps, order means nothing Addressing a specific trial

The number is what the seed derivation is a function of, never the id. Two runs of the same study with the same seed agree because trial 151 is trial 151 on both; see Determinism.

A finished trial is handed back as an immutable snapshot, which both languages call a FrozenTrial. It carries its number, its state, its parameters and the distributions they were drawn from, its objective values, the intermediate values it reported as it ran, timestamps, and — if it was a fork or part of a multi-seed fan — its parent and its per-seed results. It also carries user-owned JSON attributes. Study attributes live beside the immutable study configuration; keys under atune: are reserved for generated interoperability metadata.

The states a trial moves through

There are six, and they are the same six in both languages (Python spells them in lower case as strings):

State Meaning
Waiting Created, not yet picked up by a worker
Running A worker is evaluating it
Paused Suspended by a scheduler, holding a checkpoint reference; may resume
Complete Finished successfully; its values are the objective
Pruned Stopped early by a scheduler
Failed Stopped by an error, which is recorded on the trial

Complete, Pruned and Failed are terminal, in the strong sense that no transition leads out of them at all: the state a trial finishes in is the state it keeps, and writing onto a trial you know to be finished is refused as a conflict. Losing a race is a different thing and is not an error at all — a worker that still believed the trial was running is simply told it did not win, because two workers reaching for the same trial is ordinary rather than exceptional. Waiting and Running are the ordinary path; Running can also go back to Waiting, which is how a worker that dies mid-trial returns its work to the pool rather than losing it.

Every one of those moves is a durable record rather than a field someone overwrites. On a journal backend a state change is one transition line carrying the state it came from and the state it went to, appended only by the worker that won the change. That is what makes a trial's history auditable, and it is why replaying a journal reconstructs the study rather than approximating it.

Three rules about the terminal states decide what a study actually means, and all three are load-bearing:

  • A pruned trial's objective value is its last intermediate value. It is not discarded and it is not treated as infinitely bad. This is what makes successive-halving schedulers correct, and it is why a pruning example's learning curve must fall with the resource — a rising curve would make being pruned early look like winning. Samplers and schedulers returns to this.
  • A failed trial keeps its record and disappears from the search. The parameters, the error text and the timestamps all survive, so you can go and read what happened; but samplers do not learn from a failure, because a crash says nothing about the objective.
  • A trial that never started keeps its number. Numbers are never reused or backfilled, so the sequence stays contiguous and the seeding stays honest.

The budget

A budget is three independent optional limits, not a choice between three kinds: a maximum number of trials, a maximum total fidelity in the study's own resource unit, and a wall-clock deadline. Leave one unset and it does not apply; set several and the study stops at whichever binds first.

The trial limit is the global count of durable records in storage, not a per-handle counter. Concurrent processes cannot overshoot it, and a trial whose record was created but whose startup failed still consumed one slot. Fidelity likewise records work consumed by a fan even when its scheduler prunes it. Unlike the atomic trial-record limit, however, the fidelity limit is a cooperative ask-time stop signal read from durable reports. Concurrent in-flight trials can finish beyond it; it is not a cross-process reservation.

The deadline is compared against the study's clock, never read from the operating system directly — which is what lets a test drive a deadline-bounded study deterministically.

Reading a study back

You want Rust Python
The best trial study.best_trial()? study.best_trial
Every trial study.view()? and its filters study.trials
Every trial in a given state the view's filters study.get_trials(states=["complete"])
The Pareto front of a multi-objective study study.pareto_front()? study.pareto_front() or study.best_trials
A table for analysis the view's iterators study.trials_dataframe()

"Best trial" is deliberately absent for a multi-objective study: there is no total order on a Pareto front, and quietly ranking by the first objective would be a lie rather than a convenience. Ask for the front instead.

Note also which trials a sampler is allowed to see: completed and pruned ones, never failed ones. That is a smaller set than "every trial", and the difference matters when you are reasoning about why a sampler proposed what it did.

Where to go next

If you want to Go to
Run one, in order, from nothing Your first study
Know what a parameter can be Search spaces
Know who proposes and who prunes Samplers and schedulers
Keep a study past the process that made it Storage
Read a finished study from the shell Inspect a study
See a trial's whole life as records on disk Journal format
Look up an error a trial failed with Error reference