用扩散模型提升机器人语言控制能力,实现跨机械臂的通用操作。
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
- 引入百亿参数扩散动作专家,支持跨机器人形态学习。
- 仅靠语言指令完成折叠衣物等复杂长程任务,无需任务微调。
- 适合需要快速适配新机器人和新任务的智能控制场景。
让机器人在多样环境中完成多种任务是机器人学习的核心挑战。尽管视觉-语言-动作(VLA)模型在泛化机器人技能方面展现出潜力,但其实际效果受限于动作表示方式与训练效率。现有VLA模型多聚焦于扩大视觉-语言模型(VLM)规模,而动作空间表示仍是关键瓶颈。本文提出DexVLA框架,通过一个百亿参数的扩散动作专家,增强VLA在复杂、长程任务中的效率与泛化能力。采用新型实体课程学习策略:(1)在跨机器人形态数据上独立预训练扩散专家;(2)对特定机器人形态对齐VLA模型;(3)快速后训练以适应新任务。在单臂、双臂及灵巧手等多种机器人形态上进行实验,验证了DexVLA无需任务微调即可应对挑战性任务,能在少量数据下学习新形态的灵巧操作,并仅凭语言提示完成如洗衣折叠等复杂长程任务。在所有设置中,性能均优于Octo、OpenVLA和Diffusion Policy等前沿模型。
原文摘要 · Abstract (English)
Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning. While vision-language-action (VLA) models have shown promise for generalizable robot skills, realizing their full potential requires addressing limitations in action representation and efficient training. Current VLA models often focus on scaling the vision-language model (VLM) component, while the action space representation remains a critical bottleneck. This paper introduces DexVLA, a novel framework designed to enhance the efficiency and generalization capabilities of VLAs for complex, long-horizon tasks across diverse robot embodiments. DexVLA features a novel diffusion-based action expert, scaled to one billion parameters, designed for cross-embodiment learning. A novel embodiment curriculum learning strategy facilitates efficient training: (1) pre-training the diffusion expert that is separable from the VLA on cross-embodiment data, (2) aligning the VLA model to specific embodiments, and (3) post-training for rapid adaptation to new tasks. We conduct comprehensive experiments across multiple embodiments, including single-arm, bimanual, and dexterous hand, demonstrating DexVLA's adaptability to challenging tasks without task-specific adaptation, its ability to learn dexterous skills on novel embodiments with limited data, and its capacity to complete complex, long-horizon tasks using only direct language prompting, such as laundry folding. In all settings, our method demonstrates superior performance compared to state-of-the-art models like Octo, OpenVLA, and Diffusion Policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。