arXiv:2603.07660cs.CV2026-03被引 5

用视频自动生成大规模3D空间数据集,提升模型空间理解能力。

Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

  • 从原始视频自动构建3D空间数据,无需人工标注
  • 包含12000个优化3DGS场景和120万条空间问答对
  • 适合研究3D视觉语言模型与空间推理的学者

追求空间智能的核心在于获取大规模、细粒度的3D数据。现有方法主要依赖少量人工标注数据生成问答对,难以扩展,且因数据域偏差限制模型性能。本文提出Holi-Spatial,首个完全自动化、大规模、空间感知的多模态数据集,基于原始视频输入,通过所提数据清洗流程构建。该数据集支持多层次空间监督,包括几何精确的3D Gaussian Splatting(3DGS)重建、深度图、物体级与关系语义标注,以及对应的空间问答对。遵循系统化流程,我们进一步构建了Holi-Spatial-4M,首个高质量大规模3D语义数据集,包含12,000个优化3DGS场景、130万张2D掩码、32万3D边界框、32万实例描述、120万3D定位实例和120万空间问答对,覆盖多样几何、关系与语义推理任务。Holi-Spatial在扫描数据集ScanNet、ScanNet++和DL3DV上显著优于现有前馈与逐场景优化方法。将视觉语言模型(VLMs)在该数据集上微调后,其空间推理性能也获得显著提升。

原文摘要 · Abstract (English)

The pursuit of spatial intelligence fundamentally relies on access to large-scale, fine-grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question-answer (QA) pairs from a limited number of manually annotated datasets, rather than systematically annotating new large-scale 3D scenes from raw web data. As a result, their scalability is severely constrained, and model performance is further hindered by domain gaps inherent in these narrowly curated datasets. In this work, we propose Holi-Spatial, the first fully automated, large-scale, spatially-aware multimodal dataset, constructed from raw video inputs without human intervention, using the proposed data curation pipeline. Holi-Spatial supports multi-level spatial supervision, ranging from geometrically accurate 3D Gaussian Splatting (3DGS) reconstructions with rendered depth maps to object-level and relational semantic annotations, together with corresponding spatial Question-Answer (QA) pairs. Following a principled and systematic pipeline, we further construct Holi-Spatial-4M, the first large-scale, high-quality 3D semantic dataset, containing 12K optimized 3DGS scenes, 1.3M 2D masks, 320K 3D bounding boxes, 320K instance captions, 1.2M 3D grounding instances, and 1.2M spatial QA pairs spanning diverse geometric, relational, and semantic reasoning tasks. Holi-Spatial demonstrates exceptional performance in data curation quality, significantly outperforming existing feed-forward and per-scene optimized methods on datasets such as ScanNet, ScanNet++, and DL3DV. Furthermore, fine-tuning Vision-Language Models (VLMs) on spatial reasoning tasks using this dataset has also led to substantial improvements in model performance.

3D重建空间推理多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。