SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
Predicting future dynamics where actions are generated.
1Advanced Robotics Centre, National University of Singapore
2MMLab, The University of Hong Kong
3Nanyang Technological University
4Singapore Institute of Manufacturing Technology, Agency for Science, Technology and Research (A*STAR)
†Corresponding author




Overview

World Action Models should predict what changes, where it changes, and how those changes inform the next robot action. Existing approaches often model either expensive observation-space targets or auxiliary latents that are not fully aligned with the acting policy.
SG-WAM learns action-conditioned future dynamics inside the policy-derived representation space and grounds that space with geometric supervision. With a 0.9B model and no large-scale embodied pretraining, it reaches 98.5% on LIBERO and 73.0% on LIBERO-Plus.
Framework
SG-WAM unifies latent future prediction, geometric grounding, and action generation in the policy representation space.

Policy-derived dynamics
Learnable dynamics tokens collect policy-relevant context and model how the scene evolves under action.
Self-guided prediction
An EMA copy of the policy supplies stable future targets in the same representation family used to act.
Geometric grounding
Geometric supervision gives policy image tokens spatial structure without adding inference-time cost.
Experiment
A 0.9B SG-WAM model achieves strong in-distribution performance and robust zero-shot transfer without large-scale embodied pretraining.
Simulation Benchmarks
SG-WAM is trained once and evaluated across in-distribution and zero-shot transfer settings.
LIBERO
| Method | Params | Embodied PT. | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|---|---|
| OpenVLA-OFT | 7B | Y | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 |
| pi0 | 3.3B | Y | 98.0 | 96.8 | 94.4 | 88.4 | 94.4 |
| pi0-FAST | 3.3B | Y | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 |
| pi0.5 | 3.3B | Y | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 |
| GR00T N1.6 | 3B | Y | 97.7 | 98.5 | 97.5 | 94.4 | 97.0 |
| Spatial Forcing | 7B | Y | 99.4 | 99.6 | 98.8 | 96.0 | 98.5 |
| WorldVLA | 7B | N | 87.6 | 96.2 | 83.4 | 60.0 | 81.8 |
| LAPA | 7B | Y | 55.4 | 58.8 | 74.6 | 73.8 | 65.7 |
| RynnVLA-002 | 7B | N | 99.0 | 99.8 | 96.4 | 94.4 | 97.4 |
| Mantis | 5.8B | Y | 98.8 | 99.2 | 94.4 | 94.2 | 96.7 |
| UniVLA | 7B | Y | 96.5 | 96.8 | 95.6 | 92.0 | 95.2 |
| Fast-WAM | 6B | N | 98.2 | 100.0 | 97.0 | 95.2 | 97.6 |
| VLA-JEPA | 2B | N | 94.8 | 99.6 | 95.8 | 94.0 | 96.1 |
| SG-WAM | 0.9B | N | 99.4 | 99.8 | 98.6 | 96.2 | 98.5 |
LIBERO-Plus
| Method | Params | Camera | Robot | Language | Light | Background | Noise | Layout | Overall |
|---|---|---|---|---|---|---|---|---|---|
| WorldVLA | 7B | 0.1 | 27.9 | 41.6 | 43.7 | 17.1 | 10.9 | 38.0 | 25.0 |
| Spatial Forcing | 7B | 20.1 | 13.4 | 40.9 | 29.1 | 33.4 | 25.7 | 39.3 | 29.1 |
| Mantis | 5.8B | 15.7 | 41.8 | 45.9 | 45.1 | 28.9 | 39.2 | 62.5 | 39.8 |
| UniVLA | 7B | 4.3 | 50.3 | 71.8 | 59.1 | 80.0 | 25.3 | 34.3 | 41.5 |
| Fast-WAM | 6B | 16.4 | 44.5 | 68.9 | 78.2 | 53.7 | 37.7 | 60.7 | 50.0 |
| pi0 | 3.3B | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.9 | 53.6 |
| VLA-JEPA | 2B | 40.3 | 55.7 | 72.9 | 88.2 | 70.5 | 38.2 | 74.6 | 62.9 |
| OpenVLA-OFT | 7B | 56.4 | 31.9 | 79.5 | 88.7 | 93.3 | 75.8 | 74.2 | 69.6 |
| SG-WAM | 0.9B | 58.6 | 48.9 | 81.4 | 89.8 | 86.1 | 80.7 | 74.2 | 73.0 |
Real-world evaluation
Success rates under different visual perturbations. SG-WAM is shown in bold.
| Model | Pick and Place | Towel Folding | Toolbox Organization | ||||||
|---|---|---|---|---|---|---|---|---|---|
| In-Distribution | Background Shift | Light Change | Novel Object | In-Distribution | Background Shift | Light Change | Novel Object | In-Distribution | |
| VLA-JEPA | 35% | 20% | 25% | 20% | 20% | 10% | 10% | 15% | 20% |
| VPP | 30% | 15% | 10% | 10% | 35% | 15% | 15% | 10% | 30% |
| SG-WAM | 75% | 55% | 60% | 40% | 45% | 25% | 35% | 25% | 50% |
Geometry Attention

Experiment Videos
We evaluated SG-WAM in both simulation and real-world deployment.
LIBERO-Plus Zero-Shot Perturbations
Camera Viewpoint
Sensor Noise
Robot Initial State
Object Layout
Language Instruction
Background Texture
Light Condition
Real-World Deployment
Pick and Place - In-Distribution
Pick and Place - Background Shift
Pick and Place - Light Change
Pick and Place - Novel Object
Towel Folding - In-Distribution
Towel Folding - Background Shift
Towel Folding - Light Change
Towel Folding - Novel Object
Toolbox Organization - In-Distribution
Citation
@misc{zhao2027sgwam,
title = {SG-WAM: Self-Guided World Modeling in
Geometry-Aware Policy Space},
author = {Zhao, Ruiteng and Zhang, Zhengshen and Su, Yue
and Wang, Wenshuo and Li, Jiahui and Yang, Zhiyuan
and Tay, Francis E. H. and Ang, Jr., Marcelo H.
and Zhu, Haiyue},
year = {2027}
}