Rust-first Gymnasium · compiled to WebAssembly

Reinforcement-learning environments,
running natively in your browser.

Every simulation on this page is a gymnasium_rs environment running in WebAssembly on your machine, including the agent that learns below. No server, no video.

live environments
–
steps simulated here
–
steps / second
–

Each visit draws a new wall, and every copy on it runs its own physics.

Loading the WebAssembly environment catalog…

01 · The catalog

Six native environments, one contract.

Each card runs one environment from the Rust catalog, driven by a small hand-written controller, not a learned policy.

Try Press a card's controller button to switch it to random actions and see how hard the task is.

02 · Swarm lab

Thousands of worlds, each with its own physics.

Up to 16,384 copies of one environment, each with its own gravity, masses and lengths, all driven by one controller tuned for the defaults.

Try Press New random physics, or set Who acts to You and steer every copy yourself.

– –

03 · Robustness map

Where does the controller break?

Up to 16,384 environments on a grid, 64 × 64 by default. Each cell sets two physics parameters and scores the controller by its mean episode return: high, bright cells succeed; low, dark cells fail. The beacon marks the default physics.

Try Hover or tap a cell to watch that environment, then press Watch in the lab; in 3D terrain, drag to orbit.

mean returnworst –best –

04 · Train with Oniro

Watch an agent learn. In this tab.

Oniro, a Rust reinforcement-learning trainer built on Burn, trains a PPO agent from random weights on 16 gymnasium_rs CartPoles, in WebAssembly on your CPU.

Try Press Train. Within about a minute the status line reads SOLVED ✔︎.

Experiment

seeded · same seed, same curve
Oniro playground Open full screen ↗

Loads Oniro's 8 MB WebAssembly trainer. The rest of this page pauses while the trainer is on screen, so it gets your whole CPU.

CPU tierall algorithms

Try other algorithms

The full playground also offers SAC, TD-MPC2 and a Dreamer-style world model, and lets you set the seed, iterations and number of environments. The model-based algorithms are heavy for a browser CPU: expect a minute or more per update.

WebGPU tierchecking adapter…

Train GPU-resident

The environments and the learning both run on your graphics card, so no data has to travel back to the CPU between steps. Needs a browser with WebGPU and a real GPU.

05 · Use it

The Gymnasium API, from Rust or Python.

A preview of the pre-release API. The crates and the Python package are not published yet.

use gymnasium::{Env, envs::{CartPoleConfig, CartPoleEnv}};

let mut env = CartPoleEnv::new(CartPoleConfig::default().with_seed(42))?;
let reset = env.reset()?;
let step = env.step(CartPoleEnv::action(1)?)?;
println!("reward={} terminated={}", step.reward, step.terminated);