DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum

Adaptive RL environments towards embodied RSI.

Haoqiang Kang1Yiming Zhang2Yiyang Guo1Chuying Li1Jianzhi Shen3Tianruo Rose Xu4Xiaokang Ye1Lianhui Qin1

1 UC San Diego2 Independent3 Johns Hopkins University4 Cornell University

Demo Overview

01:45 · 1080pDownload MP4 ↓

Deliver food under real-world constraints

Maximize delivery profit while respecting time, location, energy, money, food condition, and safety constraints.

Figure 1 · Delivery task

The RL training loop

UE5 executes actions and verifies outcomes. Delivery profit and safety feedback provide the reward for online RL.

Figure 2 · RL training loop PDF ↗

Online RL improves delivery performance

Validation curves PDF ↗
+54.3%

test net income

$7.60 → $11.73
Qwen3-VL-4B before and after safety-aware RL.

Adaptive environment for embodied RSI

The environment identifies the agent’s weaknesses and adapts the next round of training tasks.

1

Evaluate

Run fixed training probes.

2

Diagnose

Identify weak capabilities.

3

Generate

Generate targeted, feasible tasks.

4

Continue RL

Update the policy and repeat.

Validation curves and training probes PDF ↗
+16.5%

test net income

Adaptive curriculum vs. uniform sampling.
Same rollout budget and test suite.

A scalable environment for embodied RL

Scale across tasks and city maps, with 2,400 training trajectories per run.

Task scaling200 → 144k

available training configurations

+47% validation income
Map scaling1 → 8

training city maps

+61% validation income
Task and map scaling PDF ↗

Citation

Paper figure

Open full resolution ↗