arXiv:2512.00226cs.CVcs.AI2025-12

用2D图像和大模型生成3D场景密集语义标注,提升视觉语言理解能力。

DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation

  • 基于多视角2D图像与多模态大模型自动生成密集描述
  • 支持物体属性、空间关系与场景上下文的问答任务
  • 适合机器人导航、AR交互等需上下文理解的应用

3D理解是实现真实世界AI辅助的关键能力。高质量数据对推动3D理解研究至关重要。现有3D场景理解数据集通常提供几何和实例级信息,但缺乏细粒度语义标注,难以支持复杂的视觉-语言任务。本文提出DenseScan,一个通过自动化流程生成多层次描述的新数据集,该流程结合多视角2D图像与多模态大语言模型(MLLMs),实现对场景元素的密集描述,完整捕捉上下文相关的细节。进一步地,我们基于场景生成高阶问题,整合物体属性、空间关系与场景上下文,拓展了语义覆盖范围。通过将几何精度与语义丰富性结合,DenseScan拓展了下游任务范畴,涵盖精细化视觉-语言导航与交互式问答。实验表明,相比传统标注流程,本方法显著提升了3D环境中物体级理解与问答性能。我们公开发布标注数据集及标注流水线,以推动机器人、增强现实等领域的研究与应用。DenseScan旨在为3D场景理解开辟新路径,助力研究者在真实复杂环境中开展更富上下文意识的分析。

原文摘要 · Abstract (English)

3D understanding is a key capability for real-world AI assistance. High-quality data plays an important role in driving the development of the 3D understanding community. Current 3D scene understanding datasets often provide geometric and instance-level information, yet they lack the rich semantic annotations necessary for nuanced visual-language tasks.In this work, we introduce DenseScan, a novel dataset with detailed multi-level descriptions generated by an automated pipeline leveraging multi-view 2D images and multimodal large language models (MLLMs). Our approach enables dense captioning of scene elements, ensuring comprehensive object-level descriptions that capture context-sensitive details. Furthermore, we extend these annotations through scenario-based question generation, producing high-level queries that integrate object properties, spatial relationships, and scene context. By coupling geometric detail with semantic richness, DenseScan broadens the range of downstream tasks, from detailed visual-language navigation to interactive question answering. Experimental results demonstrate that our method significantly enhances object-level understanding and question-answering performance in 3D environments compared to traditional annotation pipelines. We release both the annotated dataset and our annotation pipeline to facilitate future research and applications in robotics, augmented reality, and beyond. Through DenseScan, we aim to catalyze new avenues in 3D scene understanding, allowing researchers and practitioners to tackle the complexities of real-world environments with richer, more contextually aware annotations.

3D理解密集标注视觉语言多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。