arXiv:2509.25681cs.ROcs.CV2025-09被引 25

用扩散模型统一视觉、语言与动作,提升机器人任务泛化能力。

dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought

  • 通过多模态思维链整合视觉、语言与控制,统一优化感知与决策。
  • 在LIBERO上达96.4%成功率,超越离散与连续动作策略。
  • 支持真实机器人部署,可处理需多步规划的复杂任务。

视觉-语言-动作(VLA)模型正成为机器人领域的下一代范式。我们提出dVLA,一种基于扩散的VLA模型,利用多模态思维链将视觉感知、语言推理与机器人控制统一于单一系统中。dVLA在单一扩散目标下联合优化感知、语言理解与动作执行,实现更强的跨模态推理能力,并在新指令与新物体上表现更好。为提升实际部署效率,引入前缀注意力掩码与键值缓存两种加速策略,在推理阶段实现约数倍提速。我们在仿真与真实世界中评估dVLA:在LIBERO基准上达到96.4%的平均成功率,持续优于离散与连续动作策略;在真实Franka机器人上成功完成多样化任务,包括需多步规划的装箱任务,展现出稳健的现实表现。这些结果证明了统一扩散框架在高性能、实用化VLA机器人中的巨大潜力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought to unify visual perception, language reasoning, and robotic control in a single system. dVLA jointly optimizes perception, language understanding, and action under a single diffusion objective, enabling stronger cross-modal reasoning and better generalization to novel instructions and objects. For practical deployment, we mitigate inference latency by incorporating two acceleration strategies, a prefix attention mask and KV caching, yielding up to around times speedup at test-time inference. We evaluate dVLA in both simulation and the real world: on the LIBERO benchmark, it achieves state-of-the-art performance with a 96.4% average success rate, consistently surpassing both discrete and continuous action policies; on a real Franka robot, it succeeds across a diverse task suite, including a challenging bin-picking task that requires multi-step planning, demonstrating robust real-world performance. Together, these results underscore the promise of unified diffusion frameworks for practical, high-performance VLA robotics.

机器人扩散模型多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。