用大模型直接预测可动物体操作区域,省去复杂感知系统
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
- 用预训练视觉变换器+小数据微调,实现零件级操作区域分割
- 在9.9k张仿真与真实图像上训练,跨域适应效果好
- 无需复杂数据集,适合实际机器人操作部署
视觉可行动能已成为机器人领域的重要范式,聚焦于操作前识别可交互区域。传统方法依赖像素采样或点云处理来构建行为图谱,但计算开销大且难以适应多变环境。本文提出ManipGPT框架,利用大规模预训练视觉变换器(ViT)预测可动物体的最优交互区域。我们构建了一个包含9.9k张仿真与真实图像的数据集,弥合视觉仿真到现实的差距,提升实际应用性。通过在该小规模数据集上微调ViT,显著提升了零件级操作区域分割能力,将模型的上下文分割能力适配至机器人操作场景。结合阻抗自适应策略,生成零件级操作掩码,有效实现模拟与真实环境中的操作,大幅减少对复杂数据集或感知系统的依赖。
原文摘要 · Abstract (English)
Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel sampling to identify successful interaction samples or processing pointclouds for affordance mapping. However, these approaches are computationally intensive and struggle to adapt to diverse and dynamic environments. This paper introduces ManipGPT, a framework designed to predict optimal interaction areas for articulated objects using a large pre-trained vision transformer (ViT). We create a dataset of 9.9k simulated and real images to bridge the visual sim-to-real gap and enhance real-world applicability. By fine-tuning the vision transformer on this small dataset, we significantly improve part-level affordance segmentation, adapting the model's in-context segmentation capabilities to robot manipulation scenarios. This enables effective manipulation across simulated and real-world environments by generating part-level affordance masks, paired with an impedance adaptation policy, sufficiently eliminating the need for complex datasets or perception systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。