用视频扩散模型提升机器人动作决策,兼顾空间与运动一致性。
Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
- 双系统设计:慢速视觉系统+快速动作系统协同推理
- 仿真与真实场景成功率分别提升7.7%和21.7%
- 适合需要稳定抓取与复杂运动的机器人控制任务
鲁棒感知与动力学建模是现实世界机器人策略学习的基础。现有方法虽采用视频扩散模型(VDM)增强机器人策略,提升对物理世界的理解,但忽略了VDM中帧间固有的连贯运动表征。为此,我们提出Video2Act框架,通过显式整合空间与运动感知表征,高效引导机器人动作学习。基于VDM的内在表示,我们提取前景边界与帧间运动变化,同时过滤背景噪声与任务无关偏差。这些优化后的表征作为额外条件输入扩散变压器(DiT)动作头,使其能推理操作对象及运动方式。为缓解推理效率问题,我们设计异步双系统架构:将VDM作为慢速系统2,DiT头作为快速系统1,协同生成自适应动作。通过向系统1提供运动感知条件,Video2Act在低频更新下仍保持稳定操作。评估显示,Video2Act在仿真任务中成功率超越前代最优方法7.7%,真实任务中提升21.7%,并展现出强泛化能力。
原文摘要 · Abstract (English)
Robust perception and dynamics modeling are fundamental to real-world robotic policy learning. Recent methods employ video diffusion models (VDMs) to enhance robotic policies, improving their understanding and modeling of the physical world. However, existing approaches overlook the coherent and physically consistent motion representations inherently encoded across frames in VDMs. To this end, we propose Video2Act, a framework that efficiently guides robotic action learning by explicitly integrating spatial and motion-aware representations. Building on the inherent representations of VDMs, we extract foreground boundaries and inter-frame motion variations while filtering out background noise and task-irrelevant biases. These refined representations are then used as additional conditioning inputs to a diffusion transformer (DiT) action head, enabling it to reason about what to manipulate and how to move. To mitigate inference inefficiency, we propose an asynchronous dual-system design, where the VDM functions as the slow System 2 and the DiT head as the fast System 1, working collaboratively to generate adaptive actions. By providing motion-aware conditions to System 1, Video2Act maintains stable manipulation even with low-frequency updates from the VDM. For evaluation, Video2Act surpasses previous state-of-the-art VLA methods by 7.7% in simulation and 21.7% in real-world tasks in terms of average success rate, further exhibiting strong generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。