Evaluate
Run fixed training probes.
Adaptive RL environments towards embodied RSI.
1 UC San Diego2 Independent3 Johns Hopkins University4 Cornell University
Maximize delivery profit while respecting time, location, energy, money, food condition, and safety constraints.
UE5 executes actions and verifies outcomes. Delivery profit and safety feedback provide the reward for online RL.
$7.60 → $11.73
Qwen3-VL-4B before and after safety-aware RL.
The environment identifies the agent’s weaknesses and adapts the next round of training tasks.
Run fixed training probes.
Identify weak capabilities.
Generate targeted, feasible tasks.
Update the policy and repeat.
Adaptive curriculum vs. uniform sampling.
Same rollout budget and test suite.
Scale across tasks and city maps, with 2,400 training trajectories per run.
available training configurations
+47% validation incometraining city maps
+61% validation income