arXiv:2511.12436cs.RO2025-11被引 19

构建多模态交互能力数据集,提升机器人抓取与导航的精准理解。

RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation

  • 用生成式AI构建包含200万问答标注的大规模数据集
  • 在3项任务中显著提升视觉语言模型对物体与空间交互点的识别能力
  • 适合研究机器人操作、场景理解与具身智能的学者使用

机器人操作与导航是具身智能的基础能力,依赖对环境的完整理解,包括目标物体识别、物体功能属性识别以及空间可操作区域判断。尽管视觉语言模型在高层任务规划和场景理解方面表现优异,但在推断具体物理交互位置(如有效抓取点、允许放置区)方面仍存在不足,根源在于训练数据缺乏细粒度的物体与空间功能标注。为此,我们提出RoboAfford++,一个用于多模态功能学习的大规模生成式数据集,包含869,987张图像及200万条问答标注,覆盖三类关键任务:基于属性与空间关系的目标物体识别、功能部位预测、空间可操作区域定位。同时,我们构建了RoboAfford-Eval基准测试,涵盖338个精细标注样本。实验表明,现有视觉语言模型在功能理解上存在明显缺陷,而基于RoboAfford++微调后,其对物体与空间功能的推理能力显著提升,验证了该数据集的有效性。

原文摘要 · Abstract (English)

Robotic manipulation and navigation are fundamental capabilities of embodied intelligence, enabling effective robot interactions with the physical world. Achieving these capabilities requires a cohesive understanding of the environment, including object recognition to localize target objects, object affordances to identify potential interaction areas and spatial affordances to discern optimal areas for both object placement and robot movement. While Vision-Language Models (VLMs) excel at high-level task planning and scene understanding, they often struggle to infer actionable positions for physical interaction, such as functional grasping points and permissible placement regions. This limitation stems from the lack of fine-grained annotations for object and spatial affordances in their training datasets. To tackle this challenge, we introduce RoboAfford++, a generative AI-enhanced dataset for multimodal affordance learning for both robotic manipulation and navigation. Our dataset comprises 869,987 images paired with 2.0 million question answering (QA) annotations, covering three critical tasks: object affordance recognition to identify target objects based on attributes and spatial relationships, object affordance prediction to pinpoint functional parts for manipulation, and spatial affordance localization to identify free space for object placement and robot navigation. Complementing this dataset, we propose RoboAfford-Eval, a comprehensive benchmark for assessing affordance-aware prediction in real-world scenarios, featuring 338 meticulously annotated samples across the same three tasks. Extensive experimental results reveal the deficiencies of existing VLMs in affordance learning, while fine-tuning on the RoboAfford++ dataset significantly enhances their ability to reason about object and spatial affordances, validating the dataset's effectiveness.

机器人操作多模态学习具身智能视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。