arXiv:2507.23734cs.CVcs.RO2025-07ICCV被引 18

构建大规模推理型抓取感知数据集,提升机器人在开放世界中的泛化抓取能力。

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

  • 基于人类指令构建27.3万张图像的推理型抓取感知数据集
  • 模型在真实机器人任务中实现强开放世界泛化性能
  • 适合研究具身智能与视觉语言模型泛化能力的学者

通用机器人抓取系统需在多样开放世界场景中准确理解物体功能,遵循人类指令进行操作。然而现有研究受限于缺乏基于推理的大规模功能预测数据,难以保障开放世界的有效性。为此,我们构建了一个面向抓取任务的大型功能分割基准RAGNet,包含27.3万张图像、180个类别和2.6万条推理型指令。图像涵盖野外、机器人、第一人称及仿真等多种具身数据域,并配有精细的功能图标注。语言指令通过移除类别名仅保留功能描述,显著提升难度。我们还提出名为AffordanceNet的综合功能感知抓取框架,包含在大规模功能数据上预训练的视觉语言模型和以功能图为条件的抓取网络。大量实验表明,该模型在功能分割基准和真实机器人操控任务中均展现出强大开放世界泛化能力。数据与代码已开源。

原文摘要 · Abstract (English)

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-world effectiveness. To address this limitation, we build a large-scale grasping-oriented affordance segmentation benchmark with human-like instructions, named RAGNet. It contains 273k images, 180 categories, and 26k reasoning instructions. The images cover diverse embodied data domains, such as wild, robot, ego-centric, and even simulation data. They are carefully annotated with an affordance map, while the difficulty of language instructions is largely increased by removing their category name and only providing functional descriptions. Furthermore, we propose a comprehensive affordance-based grasping framework, named AffordanceNet, which consists of a VLM pre-trained on our massive affordance data and a grasping network that conditions an affordance map to grasp the target. Extensive experiments on affordance segmentation benchmarks and real-robot manipulation tasks show that our model has a powerful open-world generalization ability. Our data and code is available at https://github.com/wudongming97/AffordanceNet.

机器人抓取视觉语言模型功能感知开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。