让机器人模型在视觉变化下仍能可靠预测动作。
Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

- 用可学习查询标记对齐未来场景语义与动作流,增强鲁棒性。
- 在保持大规模视频生成预训练的同时提升对抗视觉变化的能力。
- 适合需要稳定视觉适应性的机器人控制研究者使用。
主流世界-动作模型(WAM)依赖视频生成模型(VGM)的先验动态进行动作预测,但这些模型通常在变分自编码器(VAE)潜空间中训练,侧重像素重建,导致在光照等视觉变化下动作预测脆弱。已有工作尝试在语义潜空间构建模型以提升鲁棒性,但无法利用仅存在于VAE空间中的大规模预训练。为此,我们提出Robust-WAM,一种通用的后训练方法:保留VAE生成路径,并在动作流上添加轻量级语义前瞻对齐目标。通过可学习查询标记将未来场景语义引入动作流,使其输出隐藏状态与未来真实帧的语义前瞻对齐;并为每个查询赋予对应动作标记的位置编码以建立时序对应。在分布外泛化仿真基准和真实机器人实验中,Robust-WAM显著提升多个基线模型的成功率,且不损害分布内性能。
原文摘要 · Abstract (English)
Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs are typically trained in a variational autoencoder (VAE) latent space. However, the VAE latent space is optimized for pixel reconstruction, which rewards fine appearance detail and leaves the action prediction fragile under visual shifts. Recent works build WAMs in semantic latent space, which are more robust to appearance shifts. However, these models cannot leverage the large-scale VGM pretraining that exists only in VAE space. To overcome this dilemma, we propose Robust-WAM, a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream. This retains the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics that stay reliable under illumination shifts and other visual out-of-distribution conditions. Specifically, we employ learnable query tokens to bring future-scene semantics into the action stream by aligning their output hidden states with the semantic foresight of future ground-truth frames. To establish the temporal correspondence between each query and the future step it describes, we give it the positional encoding of the matching action tokens. Experiments on out-of-distribution generalization simulation benchmarks and a real-robot setup show that our Robust-WAM consistently improves the success rates of multiple WAM baselines without sacrificing in-distribution performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。