用视觉大模型实现少样本双臂操作,高效适配新物体类别
Bi-Adapt: Few-shot Bimanual Adaptation for Novel Categories of 3D Objects via Semantic Correspondence
- 基于视觉大模型建立语义对应关系,跨类别映射操作属性
- 仅需少量数据微调,即可零样本泛化到未见物体类别
- 在仿真与真实场景中均表现优异,适合少样本工业应用
双臂协作操作对机器人完成复杂任务至关重要,但现有方法依赖大量数据采集与训练,难以高效泛化到新颖类别的未知物体。本文提出Bi-Adapt框架,通过视觉基础模型的强表征能力,实现跨类别操作属性映射。在新类别上进行受限数据微调后,该方法可零样本泛化至未见物体,在多个基准任务中均取得高成功率。大量仿真与真实环境实验验证了其有效性与高效性,显著降低数据需求。
原文摘要 · Abstract (English)
Bimanual manipulation is imperative yet challenging for robots to execute complex tasks, requiring coordinated collaboration between two arms. However, existing methods for bimanual manipulation often rely on costly data collection and training, struggling to generalize to unseen objects in novel categories efficiently. In this paper, we present Bi-Adapt, a novel framework designed for efficient generalization for bimanual manipulation via semantic correspondence. Bi-Adapt achieves cross-category affordance mapping by leveraging the strong capability of vision foundation models. Fine-tuning with restricted data on novel categories, Bi-Adapt exhibits notable generalization to out-of-category objects in a zero-shot manner. Extensive experiments conducted in both simulation and real-world environments validate the effectiveness of our approach and demonstrate its high efficiency, achieving a high success rate on different benchmark tasks across novel categories with limited data. Project website: https://biadapt-project.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。