Dual-Space Future Experts
DSFE jointly predicts future VAE latents and DINO features through a three-branch Mixture-of-Transformers. VAE latents retain fine-grained visual dynamics, while DINO provides visually stable semantic transitions.
World Action Models / Robust Robot Learning
Overview
World Action Models jointly model robot actions and future visual dynamics, but pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content. Under visual distribution shifts, predicted futures can hallucinate training-domain content rather than remain faithful to the current scene.
ST-WAM uses DINOv3 as a shared semantic representation for future prediction and history retrieval, while retaining fine-grained VAE dynamics. It is trained end-to-end, needs no additional embodied pretraining, and performs action-only inference without explicitly generating future video.
Training-Distribution Hallucination
of manually audited predictions distinctly hallucinate training-domain content under visual distribution shifts.
LingBot-VA and Fast-WAM-Joint were each evaluated on 30 randomly sampled cases under background, illumination, and camera-viewpoint shifts. Representative audited cases are shown below.
Audited Video Cases
Each video is a manually labeled hallucination case. The top panel shows the predicted future and the bottom panel shows the policy rollout.
Method
DSFE jointly predicts future VAE latents and DINO features through a three-branch Mixture-of-Transformers. VAE latents retain fine-grained visual dynamics, while DINO provides visually stable semantic transitions.
CAIR uses the current visual-language context to retrieve task-relevant evidence from recent DINO history, producing compact intent tokens that guide action generation under visual shifts.
Structured Routing
The asymmetric attention mask lets clean current VAE and DINO tokens form anchors, allows the two noisy future streams to refine each other, and keeps action queries isolated from future targets.
Experiments
Zero-Shot Generalization
Trained on LIBERO and evaluated directly across seven visual perturbations.
| Method | Emb. PT. | Camera | Robot | Lang. | Light | BG | Noise | Layout | Overall |
|---|---|---|---|---|---|---|---|---|---|
| UniVLA | Yes | 1.8 | 46.2 | 69.6 | 69.0 | 81.0 | 21.2 | 31.9 | 42.9 |
| pi0 | Yes | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.9 | 53.6 |
| pi0-FAST | Yes | 65.1 | 21.6 | 61.0 | 73.2 | 73.2 | 74.4 | 68.8 | 61.6 |
| RIPT-VLA (OFT) | Yes | 55.2 | 31.2 | 77.6 | 88.4 | 91.6 | 73.5 | 74.2 | 68.4 |
| OpenVLA-OFT | Yes | 56.4 | 31.9 | 79.5 | 88.7 | 93.3 | 75.8 | 74.2 | 69.6 |
| X-VLA | Yes | 23.4 | 89.7 | 75.7 | 88.2 | 96.0 | 62.7 | 71.8 | 71.4 |
| Fast-WAM | No | 16.4 | 44.5 | 68.9 | 78.2 | 53.7 | 37.7 | 60.7 | 51.5 |
| Fast-WAM-Joint | No | 34.0 | 55.1 | 88.9 | 90.0 | 44.9 | 33.3 | 73.4 | 59.0 |
| ST-WAM (Ours) | No | 55.4 | 60.1 | 79.3 | 93.0 | 74.2 | 79.5 | 74.3 | 72.8 |
Physical Robot
Five tasks, nominal conditions, and four unseen visual-shift settings.
| Method | Nominal Environment | Visual Distribution Shifts | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Flower | Drawer | Scoop | Fruit | Hang | Avg. | BG | Light | Obj. App. | Comp. | Avg. | |
| pi0 | 56.7 | 46.7 | 30.0 | 46.7 | 56.7 | 47.3 | 33.3 | 40.7 | 35.3 | 22.0 | 32.8 |
| Fast-WAM | 70.0 | 66.7 | 46.7 | 66.7 | 73.3 | 64.7 | 27.3 | 35.3 | 25.3 | 15.3 | 25.8 |
| w/o Semantic Future Expert | 76.7 | 73.3 | 56.7 | 73.3 | 80.0 | 72.0 | 43.3 | 50.0 | 41.3 | 29.3 | 41.0 |
| w/o CAIR | 80.0 | 76.7 | 60.0 | 76.7 | 83.3 | 75.3 | 46.0 | 52.7 | 44.0 | 32.0 | 43.7 |
| ST-WAM (Ours) | 86.7 | 80.0 | 66.7 | 76.7 | 86.7 | 79.3 | 66.0 | 70.0 | 62.0 | 48.0 | 61.5 |
Citation
@article{wang2026stwam,
title = {ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts},
author = {Wang, Mingxin and Hu, Bin and Qian, Bin and Jiang, Kaitao and Wu, Haoning and Yan, Feng and Jing, Bowen and Hao, Ruiyang and Wang, Enyi and Niu, Kangning and Yang, Yandan and Xu, Mu and Wang, Yan and Liu, Houde and Li, Tianlun},
journal = {arXiv preprint arXiv:2607.28993},
year = {2026},
url = {https://arxiv.org/abs/2607.28993}
}