用轻量空间标记加速4D世界动作建模,兼顾精度与推理速度。
WAM4D: Fast 4D World Action Model via Spatial Register Tokens

- 用空间注册标记在训练时注入几何先验,实现高效4D建模
- 在RoboTwin 2.0和真实任务上提升空间一致性,动作预测性能良好
- 适合需要快速生成精准操作策略的机器人场景
世界动作模型(WAMs)近期在联合建模未来视觉观测与可执行机器人动作方面展现出潜力。然而,多数现有WAM仍基于2D视频或潜在空间,导致视觉上合理的预测忽略3D空间约束和遮挡接触几何,难以实现精确操控。虽然几何基础模型能从视觉观测中恢复密集3D结构与运动,但强制WAM预测密集4D表示会引入昂贵的几何解码,显著拖慢因果动作生成。为解决这一权衡,本文提出WAM4D,通过轻量级空间注册标记作为训练时的未来深度读出,将预训练几何先验迁移至因果视频-动作变换器,并在推理时移除注册分支以实现轻量化。为防止非因果捷径,进一步设计因果混合注意力机制用于混合变换器(MoT)骨干网络,明确定义视频、动作与几何标记间的模态特定可见性。在RoboTwin 2.0及挑战性真实世界操纵任务上的综合实验表明,WAM4D提升了空间一致性,同时保持了高效的推理速度与竞争力的动作预测性能。
原文摘要 · Abstract (English)
World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing WAMs still operate in 2D video or latent spaces, where visually plausible rollouts miss the 3D spatial constraints and occluded contact geometry required for precise manipulation. While geometric foundation models offer strong priors for recovering dense 3D structure and motion from visual observations, forcing WAMs to predict the dense 4D representation introduces costly geometric decoding and slows down causal action generation. To address the trade-off, we present WAM4D, a fast 4D world action model that uses lightweight spatial register tokens as training-time future-depth readouts to transfer pretrained geometric priors into a causal video-action transformer, then removes the register branch for lightweight action inference. To prevent non-causal shortcuts, we further design causal mixture attention for the Mixture-of-Transformers (MoT) WAM backbone, defining modality-specific visibility among video, action, and geometry tokens. Comprehensive experiments on RoboTwin 2.0 and challenging real-world manipulation tasks show that WAM4D improves spatial consistency and achieves competitive action prediction while maintaining efficient inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。