构建首个专注3D空间推理的视觉语言数据集,破解模型依赖表面线索的难题。
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
- 设计无物体名称提示的空间查询,强制模型理解真实空间关系。
- 包含89000+人类标注查询,覆盖相对位置、视角等8类空间推理能力。
- 适合研究具身智能、机器人规划及3D视觉语言模型的学者使用。
语言与3D感知的融合对具身AI和机器人系统理解物理世界至关重要。当前3D视觉-语言研究中,空间推理能力仍被忽视,现有数据集常将语义信息(如物体名称)与空间上下文混合,导致模型依赖表面捷径而非真正理解空间关系。为此,我们提出S extsc{urprise}3D,一个专用于评估语言引导空间推理分割的新数据集。该数据集包含超过20万组视觉-语言对,覆盖900多个来自ScanNet++ v2的详细室内场景,涵盖超过2800个独特物体类别。其中89000+条空间查询由人工标注,刻意避免出现物体名称,有效缓解空间理解中的捷径偏差。这些查询全面覆盖相对位置、叙事视角、参数化视角和绝对距离推理等空间推理技能。初步基准测试显示,当前最先进的3D视觉定位方法和3D-LLMs在该任务上表现显著不足,凸显了本数据集与配套的3D空间推理分割(3D-SRS)基准套件的必要性。S extsc{urprise}3D与3D-SRS旨在推动空间感知智能的发展,为有效具身交互与机器人规划铺路。代码与数据集可在https://github.com/liziwennba/SUPRISE获取。
原文摘要 · Abstract (English)
The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a key capability for understanding spatial relationships between objects, remains underexplored in current 3D vision-language research. Existing datasets often mix semantic cues (e.g., object name) with spatial context, leading models to rely on superficial shortcuts rather than genuinely interpreting spatial relationships. To address this gap, we introduce S\textsc{urprise}3D, a novel dataset designed to evaluate language-guided spatial reasoning segmentation in complex 3D scenes. S\textsc{urprise}3D consists of more than 200k vision language pairs across 900+ detailed indoor scenes from ScanNet++ v2, including more than 2.8k unique object classes. The dataset contains 89k+ human-annotated spatial queries deliberately crafted without object name, thereby mitigating shortcut biases in spatial understanding. These queries comprehensively cover various spatial reasoning skills, such as relative position, narrative perspective, parametric perspective, and absolute distance reasoning. Initial benchmarks demonstrate significant challenges for current state-of-the-art expert 3D visual grounding methods and 3D-LLMs, underscoring the necessity of our dataset and the accompanying 3D Spatial Reasoning Segmentation (3D-SRS) benchmark suite. S\textsc{urprise}3D and 3D-SRS aim to facilitate advancements in spatially aware AI, paving the way for effective embodied interaction and robotic planning. The code and datasets can be found in https://github.com/liziwennba/SUPRISE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。