World Expert
Predict interaction-centered futures
Forecasts twelve future frames of RGB, Scalar Affordance, and Affordance Heatmap from five observed RGB frames and a language instruction.
Affordance-aware robot learning
Learning a shared interaction space where human video teaches the World—and robot experience teaches Action.
Paper overview
Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act.
We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap within the generated future World. These complementary targets ground visual prediction in task-relevant objects and interaction regions, providing a shared interaction interface across human and robot videos.
Built on a pretrained video diffusion Transformer, separately parameterized World and Action Experts use Masked Joint Self-Attention to jointly predict future RGB, Scalar Affordance, Affordance Heatmaps, and continuous actions under a unified flow-matching objective. Human videos supervise all three World streams; robot trajectories additionally supervise actions, enabling transfer without human action labels or retargeting.
Architecture
AffordanceWAM couples two separately parameterized diffusion Transformers through directional attention. The World Expert learns what changes; the Action Expert learns how the robot should act.

World Expert
Forecasts twelve future frames of RGB, Scalar Affordance, and Affordance Heatmap from five observed RGB frames and a language instruction.
Masked joint attention
Action reads RGB and Scalar Affordance. Heatmap remains auxiliary World supervision, reducing an additional RGB-rich pathway while preserving direct RGB access.
Action Expert
Generates a 12 × 7 continuous action chunk from language, proprioception, noisy actions, and evolving RGB and Scalar Affordance features.
Predicts appearance, object motion, and scene-state changes.
A scalar field identifies task-relevant objects and interaction regions without prescribing motor commands.
Rendering the field on RGB provides auxiliary supervision that encourages spatial consistency with the visual World.
World-only warmup precedes gradual Action supervision. Human clips train all three World streams; robot clips also train Action. The curriculum shifts toward joint noisy-World training and replaces clean fixed-World targets with model predictions.
Freeze the complete World path and adapt only Action using robot demonstrations without affordance labels. 75% of samples use noisy World features; 25% use completed predicted Worlds. Both paths stop gradients at the World interface.
35 World steps and 35 Action steps exchange layer-wise features during generation.
Hold the predicted World fixed for 8 Action-only refinement steps (strength 0.25), then execute the refined chunk.
Evaluation
Average success rate
+2.3 pts vs. Cosmos PolicyAverage sequence length
Best across all reported metricsMacro-average success
+32.0 pts vs. Cosmos PolicyTask-specific household manipulation with 300 demonstrations per task for AffordanceWAM.
Average success rate (%)
Demonstration budgets differ across methods; controlled ablations isolate the effects of affordance and human video.
| Method | Demos / task | Average SR (%) |
|---|---|---|
| GR00T-N1 | 300 | 49.6 |
| UVA | 50 | 50.0 |
| UWM | 1,000 | 60.8 |
| π₀ | 300 | 62.5 |
| GR00T-N1.5 | 300 | 64.1 |
| Cosmos Policy | 50 | 67.1 |
| AffordanceWAM | 300 | 69.4 |
Train on environments A, B and C; evaluate five-task sequences in unseen environment D.
| Method | 1/5 | 2/5 | 3/5 | 4/5 | 5/5 | Avg. len. |
|---|---|---|---|---|---|---|
| SuSIE | 87.0 | 69.0 | 49.0 | 38.0 | 26.0 | 2.69 |
| GR-1 | 85.4 | 71.2 | 59.6 | 49.7 | 40.1 | 3.06 |
| OpenVLA | 91.3 | 77.8 | 62.0 | 52.1 | 43.5 | 3.27 |
| CLOVER | 96.0 | 83.5 | 70.8 | 57.5 | 45.4 | 3.53 |
| UniVLA | 95.5 | 85.8 | 75.4 | 66.9 | 56.5 | 3.80 |
| π₀ | 93.8 | 85.0 | 76.7 | 68.6 | 60.1 | 3.84 |
| Seer | 94.4 | 87.2 | 79.9 | 72.2 | 64.3 | 3.98 |
| VPP | 95.3 | 88.2 | 80.3 | 72.9 | 64.5 | 4.01 |
| AffordanceWAM | 96.1 | 92.8 | 85.5 | 77.8 | 69.5 | 4.22 |
Without a shared affordance target, additional human video hurts performance. With affordance, the same data produces clear gains.
Mean ± sample SD. Data and update budgets are fixed; fusion and masking controls also match token count and trainable parameters.
| Variant | RoboCasa | CALVIN |
|---|---|---|
| Masked Joint Self-Attention | 69.4 ± 0.5 | 4.22 ± 0.01 |
| Late fusion | 65.9 ± 0.6 | 4.00 ± 0.01 |
| Independent World → Action | 64.5 ± 0.6 | 3.90 ± 0.01 |
| Open Affordance Heatmap → Action | 63.5 ± 0.6 | 3.85 ± 0.01 |
| w/o predicted-World post-training | 66.6 ± 0.6 | 4.06 ± 0.01 |
| w/o Affordance Heatmap stream | 65.4 ± 0.6 | 3.99 ± 0.01 |
Removing the Heatmap stream is a component ablation. Removing predicted-World post-training changes training exposure, not inference-time refinement.
Robot-data exposure and training updates remain fixed while affordance-annotated human video increases.
RoboCasa success rate (%)
Robot supervision held fixedSelect a point to inspect its success rate.
Vertical axis: 62–71%. Robot-data exposure and training updates are fixed.
| Human data | Success rate |
|---|---|
| 0% | 63.7 |
| 20% | 63.9 |
| 40% | 64.3 |
| 60% | 65.2 |
| 80% | 67.1 |
| 100% | 69.4 |
Eight tasks, identical 50-trajectory adaptation budgets, and 15 evaluation trials per task.
Real-world success rate (%)
50 adaptation trajectories and 15 trials per task for both methods. Fruit thresholds measure stages of the same task.
| Task / threshold | Cosmos Policy | AffordanceWAM |
|---|---|---|
| Fruit → Plate | 60.0 | 86.7 |
| Cup | 40.0 | 60.0 |
| Duck | 46.7 | 73.3 |
| Carrot | 40.0 | 86.7 |
| Spatula | 26.7 | 66.7 |
| Close the Drawer | 46.7 | 80.0 |
| Pick All the Fruit · ≥1 | 73.3 | 86.7 |
| Pick All the Fruit · ≥2 | 53.3 | 73.3 |
| Pick All the Fruit · ≥3 | 26.7 | 53.3 |
| Arrange the Flower | 6.7 | 20.0 |
Select and place the instructed fruit on the plate.
Close drawers from distinct initial configurations.
Reference
The arXiv identifier and public project links will be updated when the preprint is released.
@article{you2026affordancewam,
title = {AffordanceWAM: Affordance-Aware Joint World--Action
Modeling for Robot Manipulation},
author = {You, Jiadi and Yu, Qize and Chen, Yue and Cai, Minghong
and Zhong, Zhide and Wang, Yuran and Ping, Bowen
and Liang, Jiaqi and Shen, Zhenhao and Yan, Haodong
and Li, Yinchuan and Wu, Ruihai and Qi, Xiaojuan
and Chen, Yingcong},
journal = {arXiv preprint},
year = {2026}
}