用文本指令生成视频与动作对,提升机器人操作学习效果
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
- 双扩散模型并行生成视频与动作,保持预训练知识
- 跨模态桥接注意力提升视频与动作协同精度
- 适合需要高质量视频动作数据的机器人研究者
我们提出一种方法,从初始图像观测和机器人关节状态出发,根据文本指令生成视频-动作对。该方法自动为视频扩散模型生成动作标签,解决了动作标注缺失问题,使视频扩散模型能充分用于机器人策略学习。现有方法或采用两阶段流水线,限制跨模态信息紧密交互;或仅适配单模态扩散模型建模联合分布,难以充分利用预训练视频知识。为此,我们(1)在预训练视频扩散模型基础上,引入并行专用的动作扩散模型以保留预训练知识;(2)设计桥接注意力机制实现高效跨模态交互;(3)构建动作精炼模块,将粗粒度动作转化为低分辨率数据集下的精确控制。在多个公开基准与真实数据集上的广泛评估表明,本方法生成的视频质量更高、动作更准确,显著优于现有基线,为利用大规模视频数据进行机器人学习提供了可扩展框架。
原文摘要 · Abstract (English)
We present a method to generate video-action pairs that follow text instructions, starting from an initial image observation and the robot's joint states. Our approach automatically provides action labels for video diffusion models, overcoming the common lack of action annotations and enabling their full use for robotic policy learning. Existing methods either adopt two-stage pipelines, which limit tightly coupled cross-modal information sharing, or rely on adapting a single-modal diffusion model for a joint distribution that cannot fully leverage pretrained video knowledge. To overcome these limitations, we (1) extend a pretrained video diffusion model with a parallel, dedicated action diffusion model that preserves pretrained knowledge, (2) introduce a Bridge Attention mechanism to enable effective cross-modal interaction, and (3) design an action refinement module to convert coarse actions into precise controls for low-resolution datasets. Extensive evaluations on multiple public benchmarks and real-world datasets demonstrate that our method generates higher-quality videos, more accurate actions, and significantly outperforms existing baselines, offering a scalable framework for leveraging large-scale video data for robotic learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。