X-WAM统一实现机器人实时执行与高质量4D世界重建。
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

- 用预训练视频扩散模型预测多视角RGB-D视频,轻量级结构复用生成空间信息。
- 79.2%成功率(RoboCasa),90.7%成功率(RoboTwin 2.0),4D重建质量超越现有方法。
- 异步去噪采样提升推理效率,适合需实时交互的机器人应用。
我们提出X-WAM,一种统一的4D世界模型,将实时机器人动作执行与高保真4D世界合成(视频+3D重建)整合于单一框架中,解决了先前统一世界模型(如UWM)仅建模2D像素空间、难以平衡动作效率与建模质量的问题。为利用预训练视频扩散模型的强大视觉先验,X-WAM通过预测多视角RGB-D视频来构想未来世界,并通过轻量级结构适配——复用预训练扩散Transformer的最后几层块至专用深度预测分支——高效获取空间信息。此外,我们提出异步噪声采样(ANS),在推理时采用特定异步去噪调度:以更少步骤快速解码动作,实现高效实时执行;同时保留完整步数生成高保真视频。训练时不完全解耦时间步,而是从联合分布采样以对齐推理分布。模型在超过5,800小时机器人数据上预训练,于RoboCasa和RoboTwin 2.0基准上分别达到79.2%和90.7%的平均成功率,其生成的4D重建在视觉与几何指标上均优于现有方法。
原文摘要 · Abstract (English)
We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruction) in a single framework, addressing the critical limitations of prior unified world models (e.g., UWM) that only model 2D pixel-space and fail to balance action efficiency and world modeling quality. To leverage the strong visual priors of pretrained video diffusion models, X-WAM imagines the future world by predicting multi-view RGB-D videos, and obtains spatial information efficiently through a lightweight structural adaptation: replicating the final few blocks of the pretrained Diffusion Transformer into a dedicated depth prediction branch for the reconstruction of future spatial information. Moreover, we propose Asynchronous Noise Sampling (ANS) to jointly optimize generation quality and action decoding efficiency. ANS applies a specialized asynchronous denoising schedule during inference, which rapidly decodes actions with fewer steps to enable efficient real-time execution, while dedicating the full sequence of steps to generate high-fidelity video. Rather than entirely decoupling the timesteps during training, ANS samples from their joint distribution to align with the inference distribution. Pretrained on over 5,800 hours of robotic data, X-WAM achieves 79.2% and 90.7% average success rate on RoboCasa and RoboTwin 2.0 benchmarks, while producing high-fidelity 4D reconstruction and generation surpassing existing methods in both visual and geometric metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。