arXiv:2507.11296cs.RO2025-07ICCV被引 9

用扩散模型联合预测动作与视频,提升双臂机器人操作成功率

Diffusion-Based Imaginative Coordination for Bimanual Manipulation

  • 用多帧隐空间预测未来状态,保留关键任务特征
  • 视频预测条件于动作,推理时无需视频预测,效率更高
  • 在仿真和真实场景中成功率分别提升24.9%、32.5%

双臂操作在机器人领域至关重要,可实现工业自动化与家庭服务中的复杂任务。然而,其高维动作空间与复杂的协同需求带来了巨大挑战。尽管视频预测已被用于表征学习与控制,以捕捉丰富的动态与行为信息,但其在提升双臂协调方面的潜力尚未被充分探索。为此,我们提出一种统一的基于扩散的框架,用于视频与动作预测的联合优化。具体而言,设计了一种多帧隐空间预测策略,将未来状态编码在压缩隐空间中,保留任务相关特征。同时引入单向注意力机制,使视频预测依赖于动作,而动作预测独立于视频预测。该设计可在推理阶段省去视频预测,显著提升效率。在两个仿真基准和一个真实场景下的实验表明,相比强基线ACT方法,本方法在ALOHA上成功率提升24.9%,在RoboTwin上提升11.1%,真实实验中提升32.5%。代码与模型已公开于https://github.com/return-sleep/Diffusion_based_imaginative_Coordination。

原文摘要 · Abstract (English)

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements. While video prediction has been recently studied for representation learning and control, leveraging its ability to capture rich dynamic and behavioral information, its potential for enhancing bimanual coordination remains underexplored. To bridge this gap, we propose a unified diffusion-based framework for the joint optimization of video and action prediction. Specifically, we propose a multi-frame latent prediction strategy that encodes future states in a compressed latent space, preserving task-relevant features. Furthermore, we introduce a unidirectional attention mechanism where video prediction is conditioned on the action, while action prediction remains independent of video prediction. This design allows us to omit video prediction during inference, significantly enhancing efficiency. Experiments on two simulated benchmarks and a real-world setting demonstrate a significant improvement in the success rate over the strong baseline ACT using our method, achieving a \textbf{24.9\%} increase on ALOHA, an \textbf{11.1\%} increase on RoboTwin, and a \textbf{32.5\%} increase in real-world experiments. Our models and code are publicly available at https://github.com/return-sleep/Diffusion_based_imaginative_Coordination.

双臂操作扩散模型视频预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。