Affordance-aware robot learning

AffordanceWAM: Affordance-Aware JointWorld–Action Modeling for Robot Manipulation

Learning a shared interaction space where human video teaches the World—and robot experience teaches Action.

Jiadi You1,3,†Qize Yu2,†Yue Chen2Minghong Cai4Zhide Zhong3Yuran Wang5Bowen Ping2Jiaqi Liang2Zhenhao Shen2Haodong Yan3Yinchuan Li6Ruihai Wu2Xiaojuan Qi1,*Yingcong Chen3,*
1The University of Hong Kong2Peking University3The Hong Kong University of Science and Technology (Guangzhou)4The Chinese University of Hong Kong5National University of Singapore6Knowin AI

† Equal contribution   * Corresponding authors

arXivCode · Coming soon
01

At a glance

One interface for two embodiments

Human and robot video connected through affordance-aware World and Action modeling
A shared object-focused affordance representation bridges heterogeneous human–robot data with joint World and Action modeling.
02

Paper overview

Abstract

Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act.

We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap within the generated future World. These complementary targets ground visual prediction in task-relevant objects and interaction regions, providing a shared interaction interface across human and robot videos.

Built on a pretrained video diffusion Transformer, separately parameterized World and Action Experts use Masked Joint Self-Attention to jointly predict future RGB, Scalar Affordance, Affordance Heatmaps, and continuous actions under a unified flow-matching objective. Human videos supervise all three World streams; robot trajectories additionally supervise actions, enabling transfer without human action labels or retargeting.

Key findingAffordance turns heterogeneous human video from a source of negative transfer into scalable supervision for robot control.
03

Architecture

Joint World–Action modeling

AffordanceWAM couples two separately parameterized diffusion Transformers through directional attention. The World Expert learns what changes; the Action Expert learns how the robot should act.

World and Action Experts coupled through Masked Joint Self-Attention
Action reads RGB and Scalar Affordance features at each layer. The mask blocks Action-to-World flow and excludes Affordance Heatmap features from Action, including indirect paths through the other World streams.
W

World Expert

Predict interaction-centered futures

Forecasts twelve future frames of RGB, Scalar Affordance, and Affordance Heatmap from five observed RGB frames and a language instruction.

A

Action Expert

Generate continuous controls

Generates a 12 × 7 continuous action chunk from language, proprioception, noisy actions, and evolving RGB and Scalar Affordance features.

RGB

How the scene evolves

Predicts appearance, object motion, and scene-state changes.

Scalar Affordance

Where to interact

A scalar field identifies task-relevant objects and interaction regions without prescribing motor commands.

Affordance Heatmap

Connect regions to objects

Rendering the field on RGB provides auxiliary supervision that encourages spatial consistency with the visual World.

300+ hrsegocentric human video
≈0.4Mrobot trajectories
5 → 12observed → predicted frames
1 : 1human–robot Stage-I sampling
Stage I

Heterogeneous joint pre-training

World-only warmup precedes gradual Action supervision. Human clips train all three World streams; robot clips also train Action. The curriculum shifts toward joint noisy-World training and replaces clean fixed-World targets with model predictions.

Stage II

Task-specific post-training

Freeze the complete World path and adapt only Action using robot demonstrations without affordance labels. 75% of samples use noisy World features; 25% use completed predicted Worlds. Both paths stop gradients at the World interface.

Inference · joint generation

Generate World and Action together

35 World steps and 35 Action steps exchange layer-wise features during generation.

Inference · final refinement

Refine Action under a completed World

Hold the predicted World fixed for 8 Action-only refinement steps (strength 0.25), then execute the refined chunk.

04

Evaluation

Results across simulation and reality

RoboCasa69.4%

Average success rate

+2.3 pts vs. Cosmos Policy
CALVIN ABC→D4.22

Average sequence length

Best across all reported metrics
Real-world basic tasks74.7%

Macro-average success

+32.0 pts vs. Cosmos Policy
Simulation

RoboCasa

Task-specific household manipulation with 300 demonstrations per task for AffordanceWAM.

Average success rate (%)

GR00T-N1300 demos / task
49.6%GR00T-N1: 49.6%
UVA50 demos / task
50.0%UVA: 50.0%
UWM1,000 demos / task
60.8%UWM: 60.8%
π₀300 demos / task
62.5%π₀: 62.5%
GR00T-N1.5300 demos / task
64.1%GR00T-N1.5: 64.1%
Cosmos Policy50 demos / task
67.1%Cosmos Policy: 67.1%
AffordanceWAM300 demos / task
69.4%AffordanceWAM: 69.4%

Demonstration budgets differ across methods; controlled ablations isolate the effects of affordance and human video.

View benchmark data
RoboCasa results
MethodDemos / taskAverage SR (%)
GR00T-N130049.6
UVA5050.0
UWM1,00060.8
π₀30062.5
GR00T-N1.530064.1
Cosmos Policy5067.1
AffordanceWAM30069.4
Generalization

CALVIN ABC→D

Train on environments A, B and C; evaluate five-task sequences in unseen environment D.

CALVIN ABC→D · consecutive-task success (%) and average length
Method1/52/53/54/55/5Avg. len.
SuSIE87.069.049.038.026.02.69
GR-185.471.259.649.740.13.06
OpenVLA91.377.862.052.143.53.27
CLOVER96.083.570.857.545.43.53
UniVLA95.585.875.466.956.53.80
π₀93.885.076.768.660.13.84
Seer94.487.279.972.264.33.98
VPP95.388.280.372.964.54.01
AffordanceWAM96.192.885.577.869.54.22
Mechanism

Affordance unlocks human data

Without a shared affordance target, additional human video hurts performance. With affordance, the same data produces clear gains.

Robot only · no affordance57.3%RoboCasa3.75CALVIN
+ Human · no affordance55.1%RoboCasa ↓2.23.68CALVIN ↓0.07
Robot only · + affordance63.7%RoboCasa3.91CALVIN
+ Human · + affordance69.4%RoboCasa ↑5.74.22CALVIN ↑0.31
Architecture ablations

Why the World interface matters

Mean ± sample SD. Data and update budgets are fixed; fusion and masking controls also match token count and trainable parameters.

Architectural ablations · RoboCasa success (%) and CALVIN average length
VariantRoboCasaCALVIN
Masked Joint Self-Attention69.4 ± 0.54.22 ± 0.01
Late fusion65.9 ± 0.64.00 ± 0.01
Independent World → Action64.5 ± 0.63.90 ± 0.01
Open Affordance Heatmap → Action63.5 ± 0.63.85 ± 0.01
w/o predicted-World post-training66.6 ± 0.64.06 ± 0.01
w/o Affordance Heatmap stream65.4 ± 0.63.99 ± 0.01

Removing the Heatmap stream is a component ablation. Removing predicted-World post-training changes training exposure, not inference-time refinement.

Scaling

More human interaction experience, stronger control

Robot-data exposure and training updates remain fixed while affordance-annotated human video increases.

RoboCasa success rate (%)

Robot supervision held fixed
0% human data: 63.7% success63.720% human data: 63.9% success63.940% human data: 64.3% success64.360% human data: 65.2% success65.280% human data: 67.1% success67.1100% human data: 69.4% success69.4Affordance-annotated human data fraction

Select a point to inspect its success rate.

Vertical axis: 62–71%. Robot-data exposure and training updates are fixed.

View source data
Human-data scaling · RoboCasa success rate (%)
Human dataSuccess rate
0%63.7
20%63.9
40%64.3
60%65.2
80%67.1
100%69.4
Scaling resultThe first 60% adds 1.5 points; the final 40% adds 4.2 points—about 74% of the total gain.
Physical deployment

Real-world manipulation

Eight tasks, identical 50-trajectory adaptation budgets, and 15 evaluation trials per task.

Real-world success rate (%)

AffordanceWAMCosmos Policy
Fruit → Plate
86.7%AffordanceWAM · Fruit → Plate: 86.7%
60.0%Cosmos Policy · Fruit → Plate: 60.0%
Cup
60.0%AffordanceWAM · Cup: 60.0%
40.0%Cosmos Policy · Cup: 40.0%
Duck
73.3%AffordanceWAM · Duck: 73.3%
46.7%Cosmos Policy · Duck: 46.7%
Carrot
86.7%AffordanceWAM · Carrot: 86.7%
40.0%Cosmos Policy · Carrot: 40.0%
Spatula
66.7%AffordanceWAM · Spatula: 66.7%
26.7%Cosmos Policy · Spatula: 26.7%

50 adaptation trajectories and 15 trials per task for both methods. Fruit thresholds measure stages of the same task.

View all real-world results
Success rates (%) · all tasks and fruit-placement thresholds
Task / thresholdCosmos PolicyAffordanceWAM
Fruit → Plate60.086.7
Cup40.060.0
Duck46.773.3
Carrot40.086.7
Spatula26.766.7
Close the Drawer46.780.0
Pick All the Fruit · ≥173.386.7
Pick All the Fruit · ≥253.373.3
Pick All the Fruit · ≥326.753.3
Arrange the Flower6.720.0
Language-conditioned selection

Basic tasks

Select and place the instructed fruit on the plate.

4× · Muted
CategoryApple
Stateful and long-horizon execution

Complex tasks

Close drawers from distinct initial configurations.

4× · Muted
State transitionFirst Drawer
Fruit → plate86.7%vs. 60.0%
Close drawer80.0%vs. 46.7%
≥3 fruits53.3%vs. 26.7%
Arrange flower20.0%vs. 6.7%
05

Reference

Citation

The arXiv identifier and public project links will be updated when the preprint is released.

@article{you2026affordancewam,
  title   = {AffordanceWAM: Affordance-Aware Joint World--Action
             Modeling for Robot Manipulation},
  author  = {You, Jiadi and Yu, Qize and Chen, Yue and Cai, Minghong
             and Zhong, Zhide and Wang, Yuran and Ping, Bowen
             and Liang, Jiaqi and Shen, Zhenhao and Yan, Haodong
             and Li, Yinchuan and Wu, Ruihai and Qi, Xiaojuan
             and Chen, Yingcong},
  journal = {arXiv preprint},
  year    = {2026}
}