arXiv:2511.16857cs.CVcs.RO2025-11中稿 · CVPR被引 5

构建大规模物体交互推理数据集,提升视觉语言模型对真实场景的理解能力。

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

  • 基于BOP数据集生成6D姿态,构建细粒度物体交互标注
  • 包含150k图像与3300万问答对,覆盖六项任务(四项新任务)
  • 适合研究物体操作、空间规划与跨模态推理的学者使用

视觉语言模型在空间推理基准上表现优异,但现有评估忽略了真实应用所需的细粒度空间理解:精确3D定位、物体物理兼容性、功能属性及多步空间规划。本文提出BOP-ASK,一个大规模物体交互推理数据集,用于训练与评测。数据生成基于BOP基准中的6D物体姿态,衍生出抓取姿态、指代物体位置、路径规划轨迹、相对空间与深度关系、物间关系等细粒度标注。BOP-ASK包含超过150,000张图像和3300万组问答对,涵盖六个任务(其中四个为新任务)。我们评估了商用与开源的VLM,并在核心测试集上进行人工评估。同时发布BOP-ASK-lab,一个非BOP来源的分布外测试集,以检验泛化能力。实验表明,基于BOP-ASK训练的模型超越基线,在杂乱环境中展现出精确物体与抓取姿态估计、轨迹规划及细粒度以物体为中心的空间推理等涌现能力。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interactions. Current benchmarks test high level relationships ('left of,' 'behind', etc.) but ignore fine-grained spatial understanding needed for real world applications: precise 3D localization, physical compatibility between objects, object affordances and multi step spatial planning. In this work, we present BOP-ASK, a novel large scale dataset for object interaction reasoning for both training and benchmarking. Our data generation pipeline leverages 6D object poses from the Benchmark for Object Pose Estimation (BOP) datasets from which we derive fine grained annotations such as grasp poses, referred object poses, path planning trajectories, relative spatial and depth relationships, and object-to-object relationships. BOP-ASK comprises over 150k images and 33M question answer pairs spanning six tasks (four novel), providing a rich resource for training and evaluating VLMs. We evaluate proprietary and open sourced VLMs, and conduct human evaluations on BOP-ASK-core, a contributed test benchmark. We also release BOP-ASK-lab, an out-of-distribution benchmark with images not sourced from BOP, enabling testing of generalization. Our experiments demonstrate that models trained on BOP-ASK outperform baselines and exhibit emergent capabilities such as precise object and grasp pose estimation, trajectory planning, and fine-grained object-centric spatial reasoning in cluttered environments.

视觉语言模型物体交互空间推理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。