Autonomous disaster navigation  ·  no human entry required

When firefighters can't go in,
DEBRIS goes first.

Every year, people die trapped in burning buildings because structural collapse makes entry too dangerous for first responders. DEBRIS is a reinforcement learning agent trained to navigate smoke-filled, collapsing environments using only spatial sensors — no cameras, no visibility required. Deployable on a ground robot or drone, it finds the path to safety and guides survivors out.

untrained agent
[ video ] untrained.mp4
trained agent
[ video ] trained.mp4
start x = −12.0 m 28 m corridor exit x = +13.7 m
trained
0%
untrained
0%

How it works

Navigating blind, in the dark, under fire

A firefighter entering a burning building relies on vision, radio contact, and years of training. DEBRIS relies on physics. Its 71-dimensional sensor array — proximity sensors, velocity, spatial memory — mirrors what a person feels through vibration, touch, and spatial awareness when visibility is zero. The policy was trained entirely from physics reward signals with no human demonstrations.

01 /

No cameras required

71-dimensional sensor-only observation. Works in zero visibility, smoke, and total darkness. The same sensory modality available to a search-and-rescue robot equipped with proximity and inertial sensors — and nothing else.

02 /

Learns under active collapse

Four-stage curriculum: open corridor → falling debris → scatter shrapnel → full chaos with simultaneous ceiling crumble, fire zones, and wall collapse. Policy trained on A100 GPU via Modal, 4.3M steps with no human demonstrations.

03 /

Guides, doesn't replace

The agent finds the path. A survivor follows audio or haptic cues from the robot. Firefighters coordinate from outside. No human enters the danger zone. The robot is expendable. People are not.

Eval  ·  debris_and_scatter stage · 10 seeds

97% of corridor navigated. 1 collision. No vision.

97%
Corridor covered
trained · 10/10 seeds
1.0
Avg collisions
trained (untrained ~0.1)
4.3M
Training steps
4 curriculum stages
514
Steps / sec
A100-40GB · Modal
full_chaos stage — reward over training

Technical details

How it's built

Curriculum

Stage Debris Scatter Crumble Steps Reward
static_only 00— 600k +8.5
debris_only 30— 1.0M −2.8
debris_and_scatter 53— 1.2M +0.17
full_chaos 55✓ 1.5M −10.9

Observation space · 71-dim

IndicesNameDimEncodes
[0:3] agent_pos 3 xyz world frame
[3:9] vel + facing 6 velocity + forward vector
[9:12] goal_rel_pos 3 goal − agent
[12:47] debris pos+vz 35 10 bodies × xyz + fall speed
[47:62] lateral vel 15 vx vy of primary debris
[62:70] flags 8 time · heading · ceil · fire
[70] collision 1 contact force threshold
Physics MuJoCo 3.9 · implicitfast integrator · timestep=0.002s · velocity actuators kv=4000
Algorithm PPO · GAE λ=0.95 · γ=0.99 · clip=0.2 · 16 parallel envs · entropy coef 0.05→0.005
Policy MLP actor-critic · Linear(71→256)→LayerNorm→Tanh ×2 · orthogonal init
Infrastructure Modal A100-40GB · 72h timeout · curriculum checkpoint transfer · Modal Volume persistence