构建200万条多视角空间推理问答数据,提升模型在遮挡场景下的理解能力
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
- 设计可扩展的数据生成与标注流程,模拟真实世界部分可见场景
- 构建包含100万题目的评测基准,覆盖11类任务和多种答案格式
- 适合研究多视角视觉语言模型、空间推理与鲁棒感知的开发者使用
多模态大语言模型(MLLM)在无需显式三维重建的情况下,已能直接从多视角图像中实现3D场景理解与空间推理。然而,现实环境中常见的部分遮挡、视角重叠度低等挑战仍缺乏深入研究。为此,我们提出一种可扩展的多视角数据生成与标注管道,构建了名为SpatialMosaic的指令微调数据集,包含200万条问答对。同时,我们引入SpatialMosaic-Bench,一个涵盖11个任务、100万条问答的复杂多视角空间推理评测基准,支持选择题与数值答案。数据集覆盖室内外场景,支持多样化的现实场景评估。此外,我们通过将几何编码器集成至视觉语言模型(VLM),提供了一种实用的跨视图一致性增强基线。大量实验表明,该数据集能有效提升模型在复杂多视角条件下的空间推理能力,验证了数据生成流程在构建真实且具有挑战性问题上的有效性。
原文摘要 · Abstract (English)
Recent progress in Multimodal Large Language Models (MLLMs) has enabled 3D scene understanding and spatial reasoning directly from multi-view images, without requiring explicit 3D reconstructions. Nevertheless, key challenges that frequently arise in real-world environments, such as partial visibility, occlusion, and low-overlap conditions that require reasoning from fragmented visual cues, remain under-explored. To address these limitations, we propose a scalable multi-view data generation and annotation pipeline that constructs realistic spatial reasoning QAs, resulting in SpatialMosaic, a comprehensive instruction-tuning dataset with 2M QA pairs. We further introduce SpatialMosaic-Bench, a challenging benchmark for evaluating multi-view spatial reasoning under complex and diverse scenarios, consisting of 1M QA pairs across 11 tasks with both multiple-choice and numerical-answer formats. Our dataset spans both indoor and outdoor scenes, enabling comprehensive evaluation across diverse real-world scenarios. In addition, we provide a practical baseline for multi-view settings by integrating geometry encoders into VLMs for improved cross-view consistency and spatial grounding. Extensive experiments demonstrate that our dataset effectively enhances spatial reasoning under challenging multi-view conditions, validating the effectiveness of our data generation pipeline in constructing realistic and challenging QAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。