无需示范即可让机器人识别并操作新物体,成功率64%。
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
- 用视觉-语言-动作模型建立物体与动作的隐式关联
- 在真实机器人上对100个未见物体实现64%操作成功率
- 手机拍几张图就能微调模型,降低示范成本
模仿学习在教授机器人灵巧操作技能方面已证明非常有效,但通常依赖大量人类示范数据,限制了其在动态现实环境中的可扩展性。一个关键挑战是物体泛化:训练时针对特定物体(如“递出苹果”)掌握技能的机器人,难以迁移到语义相似但外观不同的物体(如“递出桃子”)。此前端到端视觉运动策略学习的研究尚未充分解决此类跨类别物体的泛化问题。本文提出一种基于视觉-语言-动作(VLA)模型的简单而高效方法——ObjectVLA,使机器人能在不需为每个新物体提供明确示范的情况下,泛化已有技能。通过利用视觉-语言配对数据,该方法以轻量级方式注入目标物体知识,建立物体与期望动作之间的隐式联系。我们在真实机器人平台上评估了ObjectVLA,成功在100个未见物体上实现64%的选择成功率。此外,我们提出一种更易获取的方法:仅需用手机拍摄几幅图像,即可对预训练模型进行微调,提升物体泛化能力。结果表明,该方法有效实现了物体层面的泛化,显著减少对大量人工示范的依赖,为构建更灵活、可扩展的机器人学习系统铺平道路。
原文摘要 · Abstract (English)
Imitation learning has proven to be highly effective in teaching robots dexterous manipulation skills. However, it typically relies on large amounts of human demonstration data, which limits its scalability and applicability in dynamic, real-world environments. One key challenge in this context is object generalization, where a robot trained to perform a task with one object, such as "hand over the apple," struggles to transfer its skills to a semantically similar but visually different object, such as "hand over the peach." This gap in generalization to new objects beyond those in the same category has yet to be adequately addressed in previous work on end-to-end visuomotor policy learning. In this paper, we present a simple yet effective approach for achieving object generalization through Vision-Language-Action (VLA) models, referred to as \textbf{ObjectVLA}. Our model enables robots to generalize learned skills to novel objects without requiring explicit human demonstrations for each new target object. By leveraging vision-language pair data, our method provides a lightweight and scalable way to inject knowledge about the target object, establishing an implicit link between the object and the desired action. We evaluate ObjectVLA on a real robotic platform, demonstrating its ability to generalize across 100 novel objects with a 64\% success rate in selecting objects not seen during training. Furthermore, we propose a more accessible method for enhancing object generalization in VLA models, using a smartphone to capture a few images and fine-tune the pre-trained model. These results highlight the effectiveness of our approach in enabling object-level generalization and reducing the need for extensive human demonstrations, paving the way for more flexible and scalable robotic learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。