arXiv:2512.17661cs.ROcs.LG2025-12被引 11

用视频扩散模型实现机器人闭环控制,速度快且精准。

Vidarc: Embodied Video Diffusion Model for Closed-loop Control

  • 通过掩码逆动力学模型引导视频生成,实现动作相关预测
  • 实测成功率提升15%,延迟降低91%
  • 适用于未见过的机器人平台,具备纠错和泛化能力

在数据稀缺环境下,机器人机械臂操作因复杂的具身动态和多样场景而极具挑战。现有基于视频的方法虽能通过大规模互联网视频预训练捕捉时序与物理交互,但通常不针对具身闭环控制优化,常存在高延迟和缺乏定位的问题。本文提出Vidarc(视频扩散用于动作推理与闭环控制),一种结合掩码逆动力学模型的自回归具身视频扩散方法。通过动作相关的掩码对齐视频预测,并利用缓存的自回归生成实现实时反馈,实现了快速精准的闭环控制。该模型在百万级跨具身任务数据上预训练,在真实部署中成功率达基线以上至少15%,延迟降低91%。同时展现出对未见机器人平台的强大泛化与错误纠正能力。

原文摘要 · Abstract (English)

Robotic arm manipulation in data-scarce settings is a highly challenging task due to the complex embodiment dynamics and diverse contexts. Recent video-based approaches have shown great promise in capturing and transferring the temporal and physical interactions by pre-training on Internet-scale video data. However, such methods are often not optimized for the embodiment-specific closed-loop control, typically suffering from high latency and insufficient grounding. In this paper, we present Vidarc (Video Diffusion for Action Reasoning and Closed-loop Control), a novel autoregressive embodied video diffusion approach augmented by a masked inverse dynamics model. By grounding video predictions with action-relevant masks and incorporating real-time feedback through cached autoregressive generation, Vidarc achieves fast, accurate closed-loop control. Pre-trained on one million cross-embodiment episodes, Vidarc surpasses state-of-the-art baselines, achieving at least a 15% higher success rate in real-world deployment and a 91% reduction in latency. We also highlight its robust generalization and error correction capabilities across previously unseen robotic platforms.

视频生成机器人控制扩散模型闭环系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。