用少量示范实现机器人精准抓取,无需仿真数据
ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
- 通过对象中心控制结构微调预训练视觉-语言-动作模型
- 仅需10-20次示范即达76.7%成功率,远超传统方法
- 适合低数据场景的机器人任务快速适配
学习真实世界中的机器人操作极具挑战性,尤其是在示范数据有限的情况下。现有少样本操作方法通常依赖仿真增强数据或预构建模块(如抓取、位姿估计),难以克服仿真到现实的差距且扩展性差。尽管大规模模仿预训练展现出潜力,但在数据稀缺条件下将通用策略适配到具体任务仍缺乏研究。为此,我们提出ControlVLA,一种新框架,通过类似ControlNet的结构将预训练的视觉-语言-动作(VLA)模型与对象中心表征相连接,实现高效微调。具体地,为在不覆盖先验知识的前提下引入对象中心条件,ControlVLA对一组投影层进行零初始化,使其逐步适应预训练的操作策略。在6个不同任务的真实世界实验中,包括倒置立方体和叠衣服,本方法仅需10-20次示范即达到76.7%的成功率,显著优于传统方法需超过100次示范才能达到相近效果。额外实验表明,ControlVLA可扩展至长时序任务,并对未见物体和背景具有鲁棒性。
原文摘要 · Abstract (English)
Learning real-world robotic manipulation is challenging, particularly when limited demonstrations are available. Existing methods for few-shot manipulation often rely on simulation-augmented data or pre-built modules like grasping and pose estimation, which struggle with sim-to-real gaps and lack extensibility. While large-scale imitation pre-training shows promise, adapting these general-purpose policies to specific tasks in data-scarce settings remains unexplored. To achieve this, we propose ControlVLA, a novel framework that bridges pre-trained VLA models with object-centric representations via a ControlNet-style architecture for efficient fine-tuning. Specifically, to introduce object-centric conditions without overwriting prior knowledge, ControlVLA zero-initializes a set of projection layers, allowing them to gradually adapt the pre-trained manipulation policies. In real-world experiments across 6 diverse tasks, including pouring cubes and folding clothes, our method achieves a 76.7% success rate while requiring only 10-20 demonstrations -- a significant improvement over traditional approaches that require more than 100 demonstrations to achieve comparable success. Additional experiments highlight ControlVLA's extensibility to long-horizon tasks and robustness to unseen objects and backgrounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。