arXiv:2510.27607cs.CVcs.RO2025-10中稿 · ICML被引 24

用双流扩散模型提升机器人视觉-语言-动作学习效果

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

  • 双流结构分离处理视觉与动作,通过跨模态共享知识
  • 在仿真和真实场景中分别实现6%和10%的性能提升
  • 支持无动作视频预训练,适合多源数据迁移学习

将世界模型融入视觉-语言-动作模型(VLAs)有助于机器人策略学习,但因模态差异导致状态与动作联合预测困难。为此,我们提出双流扩散框架DUST,采用多模态扩散变压器,在保持独立模态流的同时实现跨模态知识共享。DUST使用独立噪声扰动和解耦流匹配损失,学习跨模态因果关系,并引入异步采样方法,在推理时通过扩展规模提升性能。在RoboCasa和GR-1等仿真基准上,DUST相较现有最先进方法提升最高达6%,推理时扩展带来额外2%-5%增益;在Franka Research 3真实任务中,成功率提升10%。此外,DUST通过仅用无动作视频预训练及与异构机器人与人类数据联合训练,展现出良好迁移能力。

原文摘要 · Abstract (English)

Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring a multimodal diffusion transformer that maintains separate modality streams while enabling cross-modal knowledge sharing. In addition, DUST utilizes independent noise perturbations and a decoupled flow matching loss to learn cross-modal causal relationships. We further introduce an asynchronous sampling method for action and vision tokens that enhances performance through inference-time scaling. Experimental results on simulated benchmarks like RoboCasa and GR-1 show that DUST achieves up to 6% gains over state-of-the-art VLA and world-modeling baselines, with inference-time scaling providing an additional 2-5% improvement. In real-world tasks using the Franka Research 3, DUST outperforms baselines by 10% in success rate. Finally, we demonstrate that DUST enables effective transfer learning through both pretraining on action-free videos and joint-training with heterogeneous robot and human datasets.

视觉-语言-动作扩散模型机器人学习世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。