arXiv:2512.03284cs.CV2025-12被引 2

让AI像人一样主动探索3D大场景,用最少图像完成精准问答。

SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding

  • 基于文本查询主动调用空间工具,分层探索3D环境。
  • 仅需3-4张图平均完成任务,远少于基线的16+张图。
  • 适用于智能机器人、自动驾驶等需要空间理解的场景。

大规模3D场景中的空间推理对现有视觉语言模型仍具挑战,通常局限于房间级场景。我们提出H²U3D(3D房屋整体理解),一个面向房屋级场景理解的3D视觉问答数据集,涵盖最多三楼层、10-20个房间,总面积超过300平方米。通过自动化标注流程,构建层次化粗到细的视觉表征,并生成带思维链标注的多样化问答对。我们进一步提出SpatialReasoner,一种主动感知框架,能根据文本查询自主调用空间工具探索3D场景。该框架采用两阶段训练策略:先监督冷启动,再通过自适应探索奖励的强化学习,促进高效探索并抑制冗余操作。大量实验表明,SpatialReasoner在H²U3D上达到领先性能,优于GPT-4o和Gemini-2.5-Pro等强基线。尤为突出的是,本方法平均仅需3-4张图像即可完成任务,而基线需16张以上,凸显了其粗到细主动探索范式的有效性。

原文摘要 · Abstract (English)

Spatial reasoning in large-scale 3D environments remains challenging for current vision-language models, which are typically constrained to room-scale scenarios. We introduce H$^2$U3D (Holistic House Understanding in 3D), a 3D visual question answering dataset designed for house-scale scene understanding. H$^2$U3D features multi-floor environments spanning up to three floors and 10-20 rooms, covering more than 300 m$^2$. Through an automated annotation pipeline, it constructs hierarchical coarse-to-fine visual representations and generates diverse question-answer pairs with chain-of-thought annotations. We further propose SpatialReasoner, an active perception framework that autonomously invokes spatial tools to explore 3D scenes based on textual queries. SpatialReasoner is trained through a two-stage strategy: a supervised cold start followed by reinforcement learning with an adaptive exploration reward that promotes efficient exploration while discouraging redundant operations. Extensive experiments demonstrate that SpatialReasoner achieves state-of-the-art performance on H$^2$U3D, outperforming strong baselines including GPT-4o and Gemini-2.5-Pro. Notably, our method attains superior results while using only 3-4 images in total on average, compared to baselines requiring 16+ images, highlighting the effectiveness of our coarse-to-fine active exploration paradigm.

3D理解主动感知视觉问答空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。