Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Terrain Landscaping — Training Readiness

Read this first. This page records what the three landscaping tasks actually prove, and what they do not. Several readiness gates are blocked or unmeasured. No learned policy is claimed for any of them, TD-MPC2 landscaping is deliberately disabled, and no Raph hardware evidence exists. Nothing on this page should be read as a recommendation to start a long training run.

The machine-readable version of everything below is docs/superpowers/evidence/terrain_landscaping/readiness.json in this repository (it is not published on the website, and the path is deliberately not linked to a branch that may not carry it yet). Each gate cites the reviewed evidence record it reduces, in docs/superpowers/evidence/terrain_landscaping/.

1. Three separate tasks, and no curriculum

There are three landscaping task IDs. They are separate tasks with separate readiness claims, not stages of one progression.

terrain_landscaping_crater (simple)terrain_landscaping_moundterrain_landscaping (general)
Purposeone deterministic, fixed problem — the first training and deployment-test environmenta second deterministic baseline, added because the crater task’s grading metric is not reachable by this hardwarethe full task: an arbitrary hashed mission per episode
Targetthe checked-in analytic flat-bed manifest crater_target_manifest.json, hash-pinnedthe checked-in mound_target_manifest.json — the same flat analytic bed, its own reset layout mound_pile_v1a per-episode mass-balanced deformation of the flat reset bed, hashed per mission
Episode300 steps (30 s)300 steps (30 s)600 steps (60 s)
Grading cells239823986237
Resetone fixed crater + one loose pile, fixed rover poseone 0.24 m conical pile on a flat bed, no crater, fixed rover pose on the pile centrelineflat particle bed, mission drawn from a deterministic seed stream
Gate metricobservable grading MAE, margin 2.0 mmexcess volume above target, margin 1.75 mLobservable grading MAE, margin 2.0 mm

No curriculum exists anywhere in this stack. There is no automatic difficulty progression, no curriculum scheduler, no implicit switching from the crater task to the general task, and no claim that a crater-trained model solves general landscaping. The general task’s module is checked by an AST-based test that bans the identifiers curriculum, difficulty, task_crater, and the crater manifest loaders from its code, so the separation cannot silently erode.

A result on one task is not evidence for any other, in any direction.

2. What is proven, and what is not

GateSubjectStatus
G0IO contract, manifests, four hashes, no curriculumpassed
G1particle physics and readback (Isaac Sim 6.0.1 specific)passed
G2simple task is a reachable stationary-reward problemblocked — actuation
G3external-heightmap observation seampassed
G4Dreamer integration, strict config, checkpoint semanticspassed (reliability caveat resolved 2026-08-07)
G5Dreamer learnability on the simple taskblocked — measured: the model gate fails its margin
G6reusable one-environment trainingblocked — on G5 and on its own memory clause
G6-Mbatched (num_envs > 1) trainingblocked — refused fail-closed
G7TD-MPC2 on landscapingblocked by design (D13)
G8sim-to-real on real Raphblocked — no hardware evidence exists
G9general taskblocked — reachability, and no learner evidence
G2-MOUNDmound task is a reachable problem on its volume metricpassed with a caveat — a 1-in-9 per-seed flake
G4-MOUNDDreamer integration on the mound taskpassed with a scope caveat
G5-MOUNDDreamer learnability on the mound taskblocked — measured: the model gate fails its margin
G6-MOUNDreusable one-environment training, mound taskblocked — on G5-MOUND and on its own memory clause

The claim “one-environment simulation-training-ready for Dreamer” requires G0–G6. G2, G5 and G6 are blocked, so that claim is not made. The mound-task rows do not change that: G5-MOUND is blocked too, and no learned policy is claimed for any landscaping task.

G6 is blocked twice over, and unblocking G5 would not unblock it. Its randomized-profile smoke clause is gated on the G5 model gate, but its resource clause is blocked on its own measurement: peak host RSS over the canary campaign reached 23.33 GiB — 97.2 % of the declared 24 GiB limit — on a run that had only reached ~3,300 steps of a 10,000-step budget. Nothing beyond 10,000 steps was ever measured, and a declared limit may not be raised after a failing run, so the 100,000-step extension’s memory feasibility is unmeasured and on that one peak looks doubtful.

2026-08-07 re-measure: that 23.33 GiB peak came from an aborted run on the pre-fix geometry. On the fixed containment, six completed 10,000-step canaries peak at 15.9–17.6 GiB (66–73 % of the limit) with a steady-state plateau near 15.5 GiB and an essentially flat last-half slope (+0.02–0.03 GiB/h); on-disk replay is 24–26 MB against the 2 GiB limit. The projection to the 100,000-step extension is comfortable, but the clause stays blocked on its own rule: nothing beyond 10,000 steps has been measured, and the one permitted extension is deliberately unspent while G2 stands. The three seeded randomized-profile smokes have now each completed first-try (early stability evidence, recorded with a scope caveat — the plan orders them after the model gate passes).

G2 and G9 — the environment-reachability wall

Both tasks fail the same gate clause: a deterministic scripted reference must improve the final observable grading MAE by at least 2.0 mm against both a zero-action and a seeded-random baseline.

  • Simple task (G2): ten distinct scripted controller architectures were designed and measured. The best reached +0.31 mm. The retuned shipped reference lands between −0.32 mm and +0.24 mm — inside the ±0.3 mm bed-creep noise band of the zero baseline itself.
  • General task (G9): the cut-to-fill shuttle reference finishes 0.18 mm to 0.37 mm worse than doing nothing on all three declared missions.

Two caveats on the G2 number. Each of the ten architectures ran in its own OS process, and for the same simulator/reset seed the settled start state differs between processes by ~0.23 mm of observable grading MAE — the same order as the −0.49 mm..+0.31 mm spread of the campaign. Which of the ten did best is therefore not resolvable from these numbers; that they all fall ~1.5 mm short of the margin is, because 1.5 mm is far outside that effect. The per-seed rows quoted for the shipped reference are within-process and not exposed.

The evaluator is not the blocker. A teleport control that writes ~10 L of particles directly into the crater bowl moves observable MAE from 0.00915 to 0.00436 — the margin is expressible and the observable evaluator responds to it correctly. The measured mechanism is actuation: grading MAE is L1 so relocation inside the ROI is neutral, the observable map is a top-surface raster so material pushed into the bulk disappears from it, a front prismatic plate can shove but not carry, and every cut-to-fill route damages the graded surface it crosses.

Thresholds were not weakened and no episode budget was shortened. Both gate tests are committed at full strength as xfail(strict=False) so they report the deficit instead of masking it. Resolving this is an operator decision over environment-level remedies (a carrying implement, particles much finer than the blade, longer episodes, a volume-based rather than top-surface quality metric, or geometry that permits deposits without bowl transit).

G5 — blocked, and the canary commands are deliberately not published

As recorded on 2026-08-01, the 10,000-step Dreamer canary could not be completed reliably (since resolved — see the 2026-08-07 re-measure below). Of the 21 seeded canary attempt records that survived on disk for the trees under test, 1 finished its budget; 19 died with

ExternalHeightmapFrameError: mission pose [-2.00…, …, …] outside the operating envelope

killing the training process, and 1 was killed from outside. No declared resource limit was ever breached — every abort has a different cause.

The cause is measured: proprio/mission_pose is declared with low = (-2.0, -2.0, -pi), which is exactly the 4 m bed’s half-extent. The rover’s chassis centre can physically drive to and past the bed edge, so a pose 0.3 mm to 10 mm outside the bound is ordinary. The external source treats it as a frame error and raises out of env.step() rather than routing it through the episode-level safety path, so there is no truncation, no reset, and no checkpoint. The same abort is on record from A3’s scripted controller and hit 1 of 2 attempts of the 1,000-step G4 smoke, so it is not specific to the long budget — a short run is lucky, not safe.

Correction (2026-08-01). The paragraph above understates the defect, and the three remedies below were scoped against the understatement. The bed was never 4 m: PlaneCfg.size is a half-extent (spawn_plane scales a UsdGeom.Plane whose default extent is [-1, 1]², and Plane.setup_extras places the walls at ±size; srb/assets/scenery/terrain.py says so directly). Both landscaping tasks passed the declared 4 m edge length straight through, so the spawned surface was 8 m × 8 m with the 0.3 m containment walls measured at ±4.0 m — 2 m outside the D5 mission-pose envelope and outside the D6 manifest ROI. The rover was not “driving to the bed edge”; it was driving across 6 m × 6 m of bare plane that should not have existed. TaskCfg._verify_bed_fits_its_containment, the guard written to catch exactly this, computed 0.5 * min(bed_size_m) and so returned 2.0 m for a 4.0 m half-extent, which is why it passed.

The task geometry is now corrected (_plane_half_extent), and the guard measures the spawned surface instead of deriving it. No declared value changed — the D5 envelope, the D6/D8 geometry, the reset layout, every manifest hash and the io_schema_fingerprint are untouched; only the simulator now matches what they always declared. Every measured number on this page and in the A2/A3/A6a/A6b/A9 evidence records was produced against the 8 m bed and must be re-measured before it is quoted again.

Follow-up (2026-08-02). Two consequences of the same half-extent confusion were cleaned up. Plane.setup_extras derives the environment grid spacing from the plane half-extent (max(size) + 2.0), which for the corrected bed is 4.0 m — exactly the bed’s own edge length, so neighbouring environments’ containment walls would have landed in the same plane (with the old 8 m bed they overlapped by 2 m). TaskCfg.__post_init__ now pins the spacing to max(bed_size_m) + 2.0. Single-environment runs are unaffected (stack keeps env_spacing = 0.0), and num_envs > 1 remains blocked by the separate particle-transport defect. Independently, the step path no longer pulls the particle cache out of the simulator twice: the coarse LandscapingEpisodeTerrain current map it produced had no consumer, and the D7 external source already polls every environment on every step. Measured at num_envs=1: 2.0 → 1.0 simulator particle reads per step, 86.5 → 85.4 ms mean step time (median unchanged at 84.6 ms).

The one surviving checkpoint was evaluated on 2026-08-01 against all three baselines on all five predeclared evaluation seeds, in a single process per seed. It missed the gate, and was reported as indicative only (1 of the 3 required training seeds):

StatisticValueRequired
trained median final grading MAE0.009782 m
zero median0.009436 m
random median0.010276 m
scripted-reference median0.009367 m
median gain vs zero (difference of medians)−0.35 mm≥ +2.0 mm
median gain vs random (difference of medians)+0.49 mm≥ +2.0 mm
median paired per-seed difference vs zero−0.49 mm≥ +2.0 mm
median paired per-seed difference vs random−0.11 mm≥ +2.0 mm

After 10,000 steps that one trained policy was worse than doing nothing on all five evaluation seeds, and worse than the scripted reference. One training seed at the minimum budget is not evidence about learnability in either direction — which is why the full protocol was re-run once the environment defect was fixed.

Re-measure on the corrected bed (2026-08-07)

The envelope-abort defect is resolved by the containment correction above: two independent three-seed 10,000-step canary sets — six canaries — completed on the first attempt each, with zero envelope aborts, zero resource-limit breaches, zero non-finite metrics, and no retries (the reported set is exactly three seeds run once). The full model gate then executed for the first time: all three final checkpoints, evaluated against all three baselines on all five predeclared evaluation seeds, one process per seed. It fails the margin:

Statistic (3 training seeds, 15 paired rollouts)ValueRequired
trained median final grading MAE0.009943 m
zero median0.009370 m
random median0.010107 m
scripted-reference median0.009402 m
median gain vs zero (difference of medians)−0.57 mm≥ +2.0 mm
median gain vs random (difference of medians)+0.16 mm≥ +2.0 mm
median paired per-seed difference vs zero−0.42 mm≥ +2.0 mm
median paired per-seed difference vs random+0.27 mm≥ +2.0 mm

All three trained checkpoints lose to the zero-action baseline on the paired statistic (−0.47 / −0.63 / −0.06 mm per training seed), and beat seeded-random by at most +0.55 mm. Every statistic is far below the 2.0 mm margin — and the scripted reference itself achieves only +0.30 mm against zero (the G2 re-measure), so no controller, scripted or learned, can currently express the margin in a 300-step episode. G5’s blocker is therefore no longer “the canary cannot complete”: it is G2’s actuation deficit, and its remedy is the G2 remedy decision, not further training.

Consequently:

  • the conditional 100,000-step extension was still not run. Its precondition (all three canaries complete) is now met, but the plan names it the second and last qualification attempt; spending it while G2 bounds every controller below the margin would burn the one permitted attempt on a guaranteed failure. It stays unspent pending the G2 remedy decision;
  • the three randomized-profile 1,000-step smokes (a G6 clause) have now run — early, with a recorded scope caveat, since the plan orders them after the model gate passes (see the G6 note in section 2);
  • no short-canary command is documented on this page. The plan admits one only after the gate passes from a clean process; the canaries now complete reliably, but the gate itself fails, so no command is published.

Unblocking G5 was believed to require one of exactly three remedies, each reopening a different binding product decision. That framing is superseded by the correction above: remedy 3 turned out to be a plain geometry bug, and fixing it reopens nothing — the bed is now the size D6 always declared, and the rover’s own wheels stop its centre inside the D5 envelope. The 2026-08-07 re-measure confirms it: the canaries complete reliably on the corrected bed, so none of these remedies is needed any more. They are kept as the historical record of the recorded operator options:

  1. route an out-of-envelope mission_pose through the episode-level safety path (truncate and reset that environment) instead of raising out of env.step() — this reopens D11, which declares the collector step limit the only normal rollout boundary and puts real sensor/safety failures outside the MDP;
  2. widen the declared proprio/mission_pose x/y envelope beyond the terrain half-extent — this reopens D5, and therefore changes the IO-schema fingerprint, the generated RealEnv, and every checkpoint’s contract preflight;
  3. constrain the rover so its centre cannot reach the bed edge — this reopens D8, which gives the target manifest ownership of the source/ROI/guard geometry (and D6 too, but only if the source geometry itself moves; D6 merely asserts that the rover’s mission-pose centre stays inside the central 4 m ROI, which is what this remedy restores).

The mound task — reachable, and still not learned (2026-08-08)

terrain_landscaping_mound exists because of the wall above. The crater task’s gate charges the process’s own wheel and blade damage (~1 mm per active episode) against a top-surface L1 metric, and no push pattern with this blade shaves faster than it churns. The hardware is fixed — the embodiment stays the real RaphRover + RaphShovel — so the task was changed instead: level one regolith mound on a flat bed, and grade excess volume above the target rather than surface MAE. On that measure crest removal counts, ruts below target do not, and wheel churn is not charged.

The mound task is a sibling, not a mutation. The crater task, its manifest, its 2.0 mm margin, and its identity hashes are untouched. The mound carries its own reset layout (mound_pile_v1) and its own manifest hash; the D5 observation contract — the exact 1287-float actor layout and the IO-schema fingerprint — is unchanged, so both tasks speak the same interface.

The margin was frozen before the first gate run: 1.75 mL is the larger of a declared 0.5 mL floor and five times the worst measured within-process zero repeatability (0.35 mL). It has not been edited since, and it will not be. Two measured facts force the protocol that goes with it. The bed settles on its own — 5.5–7.6 mL of excess volume leaves every episode under every policy, including doing nothing. And the settled start state differs between OS processes by up to 3.6 mL for the same seed, dwarfing the ≤0.6 mL within-process spread. So each comparison runs all policies in one process from one settled start, and only within-process differences are graded.

G2-MOUND — the environment is reachable. Three consecutive full gate runs of the five-pass scripted reference passed 8 of 9 gate-seed executions (4/4, 3/4, 4/4), including two complete clean-process passes. The residual 1-in-9 flake is recorded rather than retried away: the harness retries crashes only, never assertion failures. This is the first landscaping reachability gate that passes at all.

G5-MOUND — the learner does not clear the margin. Three 10,000-step canaries completed on the first attempt each (740–771 s, 33 × 301-step episodes, no retries, no resource-limit breaches, no non-finite values), and all three checkpoints were then evaluated against both baselines and the scripted reference on all five predeclared evaluation seeds, one process per seed:

Statistic (3 training seeds, 15 paired rollouts)ValueRequired
trained median final excess volume0.018822 m³
zero median0.019368 m³
random median0.019182 m³
scripted-reference median0.016760 m³
median gain vs zero (difference of medians)+0.55 mL≥ +1.75 mL
median gain vs random (difference of medians)+0.36 mL≥ +1.75 mL
median paired per-seed difference vs zero+0.58 mL≥ +1.75 mL
median paired per-seed difference vs random+1.05 mL≥ +1.75 mL

Per training seed, the paired difference vs zero is −0.32 / +1.13 / +1.39 mL: two of the three checkpoints beat both baselines, one is worse than doing nothing, and the scripted reference still leads the trained median by 2.06 mL.

This is a different failure from the crater’s. There, the environment could not express the margin at all — the scripted reference reached +0.30 mm against a 2.0 mm requirement, bounding every controller, learned or scripted. Here the same reference clears the margin, so the environment is demonstrably reachable and what falls short is the learner at a 10,000-step budget. That is a learnability result, not an actuation wall — and it is still a failed gate. No learned mound policy is claimed.

The 100,000-step extension was spent, and it made things worse

On operator instruction the same day, all three seed runs were resumed in place to a total of 100,000 steps — the second and last qualification attempt the plan permits — and re-evaluated under the identical protocol against the identical 1.75 mL margin. The trainings were clean (332 × 301-step episodes each, 94–103 min, 8.1–8.2 GiB peak against the 24 GiB limit, 244–246 MiB replay against the 2 GiB limit, zero breaches, zero retries). The result was not:

Statistic10,000 steps100,000 stepsRequired
median gain vs zero+0.55 mL−0.09 mL≥ +1.75 mL
median gain vs random+0.36 mL−0.01 mL≥ +1.75 mL
median paired difference vs zero+0.58 mL−0.08 mL≥ +1.75 mL
median paired difference vs random+1.05 mL−0.21 mL≥ +1.75 mL
gap to the scripted reference+2.06 mL+3.20 mL

Per training seed the paired difference vs zero went −0.32 / +1.13 / +1.39 mL → −0.00 / −0.30 / +1.38 mL: only one checkpoint held its gain, and two now sit at or below doing nothing. A 10× budget did not close the gap — it erased the lead, so the 10,000-step positives read as run-to-run spread rather than an early learning trend. The “needs more steps” hypothesis is measured and refuted for this configuration, and no further extension is permitted.

One measurement is not understood and is flagged rather than smoothed over: the 100,000-step runs peak lower than the 10,000-step canaries (8.1–8.2 GiB against 17.5–17.6 GiB single-process), which is backwards for the longer run with the larger replay. No cause was established.

Steps are not the only budget axis: the replay ratio

Everything above measures the steps axis. The gradient-update budget is a separate one. Dreamer does steps × replay_ratio ÷ (batch_size × batch_length) updates, which for this stack’s batch_size: 8 and batch_length: 32 is steps × ratio ÷ 256. Both mound attempts ran at a replay ratio of 16.0 — half this repository’s own default and 1/32 of the 512 the same defaults file gives excavation, the sibling particle-manipulation task. Nothing recorded why. So the two failed attempts did roughly 625 and 6,084 gradient updates, and the diagnosis in docs/superpowers/evidence/terrain_landscaping/2026-08-12-learner-budget-diagnosis.md measured the consequence directly: the actor sat pinned at maximum entropy, wall clock was nearly all simulation, and the gate therefore evaluated an exploration distribution rather than a learned policy.

Both task profiles now declare 512.0. This changes no gate row and no result on this page. The published attempts were run at 16.0 and stay recorded as run; the raise has not itself been carried through a gate reduction at the canary budget, so nothing here is superseded and no learned policy is claimed. It also means the two failed attempts do not bound what this configuration can do — “more steps” was measured and refuted, “more updates” was never the thing under test.

Consequently:

  • the conditional 100,000-step extension is spent and failed. Its precondition (all three canaries complete) was met, and the plan permits no second one — so any next attempt is an operator decision about what changes, not another run of the same thing;
  • the three seeded randomized-profile 1,000-step mound smokes have run and were stable (three full episodes each, zero aborts, zero breaches), early and with the same scope caveat as the crater’s;
  • no short-canary command is documented on this page for the mound task either, for the same reason: the gate does not pass.

3. The actor observation contract

Both tasks publish the same seven leaves and the same three actions. The learner adapter flattens the seven leaves into exactly 1287 float32 in this order, and nothing else enters replay or the world model:

SliceLeafShapePhysical bound
[0:256]proprio_dyn/heightmap_current_global(16, 16)[-0.30, 0.50] m
[256:512]proprio_dyn/heightmap_target_global(16, 16)[-0.30, 0.50] m
[512:896]proprio_dyn/heightmap_current_local(24, 16)[-0.30, 0.50] m
[896:1280]proprio_dyn/heightmap_target_local(24, 16)[-0.30, 0.50] m
[1280:1283]proprio/mission_pose(3,)x, y ∈ [-2.0, 2.0] m, yaw ∈ [-π, π]
[1283:1286]proprio/base_velocity(3,)vx, vy ∈ [-1.5, 1.5] m/s, wz ∈ [-4.0, 4.0] rad/s
[1286:1287]proprio_dyn/heightmap_age_s(1,)[0.0, 0.5] s

Elevations are metres relative to the mission datum (the containment base plane, z = 0). The global maps are an area-mean downsample of the 4 m ROI; the local maps sample native cell centres in the rover body frame (+x forward, +y left) at the frame’s own synchronized pose, with bilinear interpolation and no extrapolation.

The actor does not receive a cell-validity mask, shovel extension, the previous action, episode time remaining, particle positions, the simulator world pose, or any privileged current-minus-target map. Values outside the bounds above are contract failures, not values to clip.

The three actions are normalized to [-1, 1]: robot/cmd_vel on [0:2] (linear scale/saturation 0.4 m/s; angular scale radians(60) then saturation at 1.0 rad/s) and payload/joint_vel on [2:3] (scale/saturation 0.04 m/s).

Full coverage is a hard invariant

There is no validity mask, so every delivered frame must be complete. The source grid is fixed: 0.05 m cells, an 80 × 80 central target ROI, a 33-cell (1.65 m) guard on every side, and therefore a 146 × 146 complete source. A frame that is non-finite anywhere, incomplete, wrongly shaped, or outside the declared elevation/pose/velocity envelope is rejected whole — there is no sentinel fill, no silent clip, and no partial frame.

Unknown cells are never turned into plausible zero elevation. A cell with no supported particle receives the calibrated physical base-plane elevation, because the source is a semantic terrain-work-surface layer (containment base plane plus regolith) and excludes the rover, shovel, containment wall, and other transient occluders. If a real lab mapper cannot provide that coverage under the rover and shovel, the full-coverage decision must be reopened rather than approximated.

4. The external-heightmap source: timing, freshness, and the target manifest

The rover has no onboard depth camera. The actor sees only the product of one external mapper.

  • 10 Hz contract. One accepted-or-held frame is published at every policy step.
  • Synchronized pose and velocity. The map, the mission pose, and the body velocity are captured in one provider snapshot and share one source timestamp. The local crop is sampled at that frame’s pose — an older map is never combined with a newer pose.
  • Source time. In simulation the declared clock is monotonic_sim_time (common_step_counter * step_dt), monotonic across episode resets. Producers reject future and non-monotonic timestamps.
  • Reset readiness. Reset blocks until the first complete frame exists, with a bounded 2.0 s timeout (20 polls at 10 Hz). Failure raises ExternalHeightmapReadinessError listing the failing environment ids. No zero startup frame is ever substituted.
  • Stale abort. The hard cutoff is 0.5 s, which is why the actor’s freshness bound is [0.0, 0.5] s. A frame older than that is rejected and handled as an out-of-MDP pause/abort — ExternalHeightmapStaleError in simulation, and a sensor_stale session abort on the deployment path.
  • Dropouts hold, they do not hole. A dropped update re-publishes the previous complete map and increases the freshness value.

Measurement profiles

Selected with env.external_heightmap.profile. Each profile is hashed (canonical JSON, SHA-256), and the hash travels in the run manifest:

ProfileBehaviourProfile hash
ideal (default)zero delay, dropout, elevation noise, registration error, pose jitter9c67bdce…2283398
randomizedwhole-frame delay U[0.0, 0.1] s; frame-drop probability 0.05 capped at two consecutive drops; per-cell elevation noise N(0, 0.002 m) clipped to ±0.006 m; one per-episode SE(2) registration translation N(0, 0.005 m) clipped to ±0.015 m and yaw N(0, 0.25°) clipped to ±0.75°; per-frame pose jitter with the same boundsa008c11d…f94870fe
failure_staleat least six consecutive held updates — used only to prove the stale failure path192fbd06…3652a8e35

These are seeded robustness defaults, not claims about lab error distributions and not a curriculum. Lab recordings may replace the randomized values only through a new versioned profile hash.

Measured agreement between the oracle (particle truth) and the observable (mapper product) evaluation, over 100 seeded frames on the crater grading cells: ideal agrees to 1e-9 with identical success classification; randomized has p95 disagreement 0.00189 m against a 0.01 m bound and 100 % success-classification agreement against a ≥ 95 % bound.

Only ideal has ever been exercised on a training run or on the general task.

The target manifest owns the mission identity

The manifest — not the code, not the config — owns the mission frame ID, the world-to-mission calibration version, the elevation datum, the source/ROI/guard geometry, the surface aggregation rule, the target map, the reset layout, the physical bounds, the grading-cell set, and the success parameters. Simulation and deployment both reject a mismatched frame, datum, geometry, or hash before policy inference.

Four hashes, never conflated. They are four different identities and conflating any two hides real drift:

HashWhat it identifiesCrater value
io_schema_fingerprintthe action/observation schema itself74bfe641…60ff16
target_map_sha256the desired elevation fieldf20f228f…f305bc8a
reset_layout_sha256the deterministic particle spawn layout6e809eb1…fd6769d4a
manifest_sha256the whole manifest document08e66df6…f1ec6f17

The crater task’s manifest is checked in and fixed; configuration fails with an actionable error if any crater/pile/spawner field stops reproducing the hashed reset layout. The general task generates a manifest per episode from a deterministic (env.general_mission_seed, env_id, episode_index) stream — so a recorded triple regenerates the identical mission in any process — and its reset_layout_sha256 (b2b2cdbf…53a95b90) is fixed for the task while its target-map and manifest hashes vary per mission.

5. Particles: GPU solver, CPU-facing readback

Landscaping regolith is a PhysX PBD particle set, and the split between where it is simulated and where it is read matters operationally:

  • The solver is GPU-only. PhysX rejects particle sets outright when GPU dynamics are unavailable, logging Particles feature is only supported on GPU. Please enable GPU dynamics flag in Property/Scene of physics scene! and leaving every particle bit-identically inert. A CUDA-capable NVIDIA GPU is required.
  • The per-particle readback is the CPU-facing USD transport (UsdGeom.Points), fed by that GPU solver. On the pinned Isaac Sim 6.0.1 build there is no direct/Fabric per-particle alternative: omni.physics.tensors.SimulationView exposes only cloth, material and system-level particle views, and isaacsim.core.prims.ParticleSystem is system-level. This was probed and recorded, not assumed.
  • Both landscaping tasks therefore set sim.device = "cpu", and SRB disables Fabric automatically whenever particles are enabled. Isaac Sim 6’s CUDA direct-data pipeline does not synchronize PBD particle positions or velocities back to the USD points that the heightmap, the reward, and the renderer consume.

One trap worth knowing. A Kit-persisted app-global setting (/persistent/physics/overrideGPUSettings = 0, “Force CPU”) overrides the authored per-scene physxScene:enableGPUDynamics=true, so a machine that once had Force-CPU selected in the UI will silently produce a completely inert particle bed while everything else — rigid bodies included — behaves normally. This was the demonstrated cause of a long-standing zero-displacement symptom; neither the particle simulationOwner relationship nor USD readback synchronization had anything to do with it. SRB now clears that override at spawn time (with a warning) rather than trusting machine state, and authors an explicit simulationOwner that must equal the configured physics_prim_path.

Every landscaping particle claim is version-specific to Isaac Sim 6.0.1 / omni.physx 110.0.7 and must be re-proven after either changes, via tests/integration/test_particle_simulation_owner.py plus the two smoke nodes.

6. Dreamer

DreamerV3 is the only learner wired to the landscaping contract. It is pinned to 4049794d4135e41c691f18da38a9af7541b01553 with elements 3.22.0.

The task-level configuration lives at hyperparams/task/terrain_landscaping_crater/dreamerv3.yaml and is merged strictly: a dropped key is fatal, and the loaded config must declare every key the checked-in file declares. An operator profile that was never written against the strict upstream schema falls back to a lenient merge with a warning, and the run records strict_task_config so you can tell which happened.

num_envs > 1 is refused at wrapper construction, before any reset, step, or allocation, with an error naming the missing selective-reset transport and the qualified env.num_envs=1 path. See §7.

Checkpoint operations — three of them, named precisely

OperationWhat it restoresWhat it does not restore
Same-logdir training-state resume (--continue)step, the agent (parameters plus optimizer/update counters), and the replay buffer, in a fresh processsimulator/environment state, partial episodes, driver and RSSM carries, process RNG
Agent import / transferagent state only, into a new run with a new logdir, starting at step 0everything else — it is never described as a resume
Evaluation / policy loadagent state only, for inferenceeverything else

The resume is not a bitwise continuation, and the exclusions above are the explicit contract, not an oversight. It requires synchronous, complete replay chunks: the run writes a ckpt_generation.json manifest recording the step, the replay item count, and the completed chunk paths, and a resume whose recorded chunks are missing or incomplete fails loudly rather than continuing on a silently empty replay. run.from_checkpoint is never used for this operation, and the process logs which operation it performed.

A portable model artifact for Dreamer is a sanitized directory, and it is not a run snapshot. It contains only the empty done marker plus exactly one agent payload form (agent.pkl xor a contiguous agent-NNNN.pkl shard set). Step and replay payloads, symlinks, nested directories and unrecognized members are structurally rejected without unpickling. Loading one starts a new run; it never resumes training.

Every Dreamer checkpoint and policy load runs a fail-closed contract preflight: the artifact’s recorded task id, IO-schema fingerprint, projection layout fingerprint, target-manifest hash, mapper-profile hash and normalization policy must all match the live environment, or the load is rejected before inference. The projection layout fingerprint is the order half — a checkpoint trained under a different packed-vector order is rejected even when its IO-schema fingerprint is identical. The only operator opt-out (allow_unverified=True / SRB_ALLOW_UNVERIFIED_LANDSCAPING_CHECKPOINT) tolerates an unverifiable artifact; it never tolerates an actual mismatch.

What Dreamer’s integration does and does not prove

G4 proves runtime viability of the one-environment strict path: the contract projection, strict configuration, upstream construct/update, truncation that resets collection while staying non-terminal for bootstrap, a 1,000-step smoke at ~14 env steps/s, checkpoint/reload/eval, the sanitized artifact path, and mismatch rejection.

It proves nothing about learnability, the general task, or hardware. Its former reliability caveat — the 1,000-step smoke aborted on the mission-pose envelope in 1 of 2 attempts of its 2026-08-01 re-measurement — is resolved (2026-08-07): on the corrected bed, five 1,000-step smokes (two ideal-profile, three randomized-profile) and six 10,000-step canaries all completed on the first attempt with zero envelope aborts.

It also does not establish that a long run fits. The G6 clause “projected replay/model memory fits the declared host envelope” is blocked on its own evidence, independently of G5: the 1,000-step smoke ends at 8.28 GB host RSS, but the peak over the 10,000-step canary campaign is 23.33 GiB — 97.2 % of the declared 24 GiB limit — at only ~3,300 steps. The limit was never breached and may not be raised after a failing run. The 2026-08-07 completed canaries re-evidence the clause (peaks at 66–73 % with a ~15.5 GiB plateau — see section 2), but nothing beyond 10,000 steps has been measured, so it stays blocked.

7. Batched training is unsupported

num_envs > 1 is refused fail-closed for Dreamer landscaping, and there are two independent measured reasons:

  1. Mixed-reset transport. After one row truncates, the current transport advances the freshly reset row by one hidden physics step under a masked action before labelling that frame is_first=True. The continuing row is otherwise unaffected and no global reset occurs, but the reset row’s first observation is not its true unstepped reset observation.
  2. Particle clone frames. ParticleSystem reads and writes per-particle state in prim-local coordinates while exposing it as world-frame buffers. With cloned environments carrying non-zero origin transforms, both beds collapse onto the world origin seam and both rover articulations read all-NaN. This is a core particle-transport defect, not a landscaping one.

Rejecting construction is the recorded outcome, deliberately, rather than weakening the mixed-reset requirement.

8. TD-MPC2 — landscaping is blocked

TD-MPC2 cannot be used for any landscaping task, and asking for it fails immediately with:

TD-MPC2 landscaping is disabled: velocity-controlled shovel extension is unobserved and no approved recurrent/history state contract exists.

The reason is observability, not plumbing. The shovel is velocity-controlled and its extension is not in the actor observation, so the same visible state can correspond to different shovel extensions. Upstream TD-MPC2 encodes the current state feed-forward and cannot recover that hidden actuator state (_prev_mean is planner state, not recurrent observation state). Dreamer is allowed a belief about extension only because its recurrent carry includes the previous action, and even that is an accepted partial-observability risk for a short canary rather than proof that extension is observed.

The block is capability/contract based, not name based — it fires on either frozen task ID in any spelling, on any environment publishing landscaping_contract_metadata(), and on any generated RealEnv declaring the contract’s IO-schema fingerprint. It fires before the log directory is created, before config.yaml is written, before checkpoint discovery, before replay allocation, and before the agent is constructed. Unblocking requires separate approval of one of: measured shovel extension, an approved observation/action history or recurrent encoder, or another physically grounded state estimator. Integrating commanded velocity and presenting the estimate as measured extension is explicitly forbidden.

What TD-MPC2 does work on

The generic dict-flattening path is repaired and usable on ordinary multi-leaf SRB tasks: every actor-visible numeric non-image leaf (packed vectors and 2-D map leaves) is concatenated into one float32 state vector with physical bounds, and image-like leaves are rejected at construction with an actionable message instead of reaching upstream’s encoder after the log directory, config, replay and model have already been created.

That path is gated on a pinned upstream checkout:

ItemValue
upstream base8bbc14ebabdb32ea7ada5c801dc525d0dc73bafe
backport appliede9f59321933cbc8e11a002b842adc7d4ffae8ff1 (fix Q-ensemble weight init)
resulting pinned SHAbfb0029669f7242c33b1400950f97518ba46a5d8
deliberately rejected75212c3a090115df212402ac911df446ccc2047f — it pins Torch 2.7.1 / TensorDict 0.8.3 / TorchRL 0.8.1, older than Isaac Sim ships

The decision was to keep Isaac’s installed Torch/TensorDict/TorchRL matrix and backport only the Q-ensemble initialization fix. Without the backport the ensemble silently keeps PyTorch’s kaiming_uniform_ weights and non-zero biases instead of the intended trunc_normal_(std=0.02) with zeroed biases. assert_upstream_compatible() runs in both the train and policy-load paths before anything is allocated and refuses an unpatched checkout with the exact re-pin commands. The upstream contract test therefore checks the installed checkout, so its result is environment-dependent: it is green against the pinned patched checkout and red against an unpatched one, which is the gate working as designed.

TD-MPC2 checkpoints are weights-only

Upstream’s save() writes {"model": state_dict} and nothing else — no optimizer, replay, step, or RNG state. Therefore:

  • --model <ckpt> on train is a weights import. Training starts at an explicitly logged step 0 with a fresh replay buffer, fresh optimizers and reset schedules. The step is not parsed out of the file name.
  • --continue / continue_training=True fails immediately with a NotImplementedError explaining that no versioned full trainer-state checkpoint exists; nothing is allocated first.
  • an evaluation/policy load is logged as Evaluation load (weights only).

No TD-MPC2 weights-only load is a training resume, and no output describes one as such. Contrast this with Dreamer, which does have a real same-logdir training-state resume (§6).

9. Real Raph validation is unavailable

No hardware evidence exists. The hardware-free half of the sim-to-real work is complete — the observable task evaluator, the session-abort vocabulary, the horizon-truncation semantics, the model-artifact refusal, and a crater validation spec that resolves its real environment. The live half was never started: no live Raph or mapper contract has been recorded, and there is no lab run.

Constructing the canonical landscaping RealEnv validates four capability tags before it acquires a ROS node, starts hardware, or lets a caller load a policy:

TagClaimed by shipped code?
landscaping.observable_task_evaluatoryes — LandscapingRealEvaluator
raph.drive_velocityno
raph.shovel_prismatic_velocityno
landscaping.external_heightmap_batchno

A CLI validation run of the crater spec therefore exits 4 with a typed DeploymentNotReadyError naming the missing tags. That is the intended and measured state.

What must be recorded from the lab before this can change: the Raph ROS drive/shovel command topics and message types, unit and sign conventions, saturation behaviour, watchdog/timeout semantics, acknowledgement mechanism, measured joint-state feedback, the measured real drive limits (the values in the contract are simulator values, not verified lab facts), the live external mapper and its calibration, live map/pose and map/velocity skew, and the live 10 Hz pacing and latency distribution.

Simulation and fake-adapter results are not hardware proof. Passing every other gate on this page would still not imply G8.

Known reporting hole. srb real_agent validate --dry-run skips environment and policy instantiation and finalizes with exit 0, pass_overall: true, n_episodes: 0, and a green real-eval: pass badge.json. Its three criteria are all reported with status skipped, which is the only in-band signal that nothing was validated. Do not treat a dry-run artifact as a validation result. Recorded, not fixed.

10. Generating and checking the deployment bridge

Both task IDs generate distinct bridge modules, and the general task must never be relabelled with the crater’s hash:

# Regenerate the checked-in bridge modules (writes in place)
srb real_agent gen --env terrain_landscaping_crater
srb real_agent gen --env terrain_landscaping

# Freshness gate: read-only byte comparison against the checked-in module
srb real_agent gen --env terrain_landscaping_crater --check
srb real_agent gen --env terrain_landscaping --check

Both run unqualified — no env.particles_height override belongs in them, and none must be reintroduced. --check renders and formats a candidate, byte-compares it against the checked-in module, prints a unified diff, and exits non-zero on drift without touching the file. It requires the repo’s pinned formatter (ruff) on PATH; without one it refuses to compare rather than manufacturing drift from an unformatted candidate. srb real_agent gen --env ALL --check runs the gate across every cached environment and exits non-zero if any drifted or if the cache is empty.

See srb real_agent and the Sim-to-Real workflow §5.2 for the generated schema and the deployment capability gate.

11. Configuration hazard worth knowing

srb.utils.hydra.extract.extract_defaults_from_class serializes an asset-instance field default as {"name": …} plus a small allow-list (action_mode and nested sub-asset fields). Any other customization the task declared on that instance is silently dropped, so a CLI/Hydra-launched run rebuilds the asset from its class defaults while a directly constructed TaskCfg(...) gets the declared values.

This was measured on landscaping: direct construction received a (4.0, 4.0) m containment with 0.3 m walls; the CLI path received Plane’s class defaults (0.75, 2.0) m with 1.0 m walls. Every CLI-launched landscaping run before the fix — crater included — trained on a surface a fraction of the declared size.

The landscaping tasks are repaired locally by declaring the containment as plain scalars (env.bed_size_m, env.bed_wall_height_m) that survive the round trip. The underlying allow-list is unchanged, so this remains a repo-wide hazard for any task that customizes an asset-instance default, and one cosmetic instance is still live in landscaping (the crater’s plane visual_material). If you customize an asset instance in a task config, verify it survives the Hydra round trip rather than assuming it does.

12. Evidence

Every claim on this page reduces one of the reviewed records in docs/superpowers/evidence/terrain_landscaping/:

RecordSubject
A0.mdfrozen IO contract, target manifest, four hashes
A1.mdparticle dynamics and readback, demonstrated cause
A2.mdcloned Raph locomotion, D4 fingerprint sync, G6-M blocker
A3.mdsimple-task reward semantics, feasibility, G2 reachability campaign
A4.mdexternal-heightmap observation seam
A5.mdsimulator / generated RealEnv / hardware schema alignment
A6a.mdDreamer integration, strict config, checkpoint semantics (G4)
A6b.mdbounded canaries and the G5 model gate (blocked)
A7.mdTD-MPC2 generic repair, upstream pin, landscaping preflight
A8a.mdtruthful real validation, hardware-free half
A10.mdgeneral task repair, feasibility, G9 reachability
A9.mdthis reduction
2026-08-07-recheck.mdtraining-path recheck, the corrected-bed re-measure
2026-08-07-mound-spike.mdmound-geometry feasibility spike, volume-metric separation
2026-08-07-mound-margin-derivation.mdthe mound volume margin and the paired protocol
2026-08-08-mound-ladder.mdthe mound-scoped ladder and its model-gate failure
readiness.jsonthe machine-readable gate report

A8b.md does not exist: the live hardware half was never started.