构建跨视角空间智能三件套,提升多模态大模型的多视角理解能力。
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark

- 设计多智能体数据引擎,生成160万样本的跨视角指令数据集
- 提出跨视角对齐框架,使模型在多视角下保持物体一致性
- 建立场景独立评测基准,系统评估模型跨视角推理能力
空间智能要求多模态大语言模型(MLLMs)超越单视角感知,一致地推理物体、可见性、几何关系及交互行为在多个视角下的表现。然而,当前跨视角推理进展受限于三大瓶颈:大规模高质量标注数据稀缺、系统性评测基准缺失、缺乏显式对象级跨视图对齐机制。为此,我们构建了由三个协同组件组成的CrossView Suite:CrossViewSet、CrossViewBench与CrossViewer。首先,通过多智能体数据引擎精心构建大规模高质跨视角指令数据集CrossViewSet,覆盖17种细粒度任务类型,共含160万样本。其次,构建场景无关的CrossViewBench,全面评估MLLM在跨视角空间理解方面的表现。最后,提出CrossViewer,一种遵循感知-对齐-推理范式的渐进式三阶段框架,采用自适应空间区域分词器捕捉细粒度物体表征,显式对齐多视角物体,并融合对齐特征以增强跨视角推理能力。大量实验与分析表明,大规模训练数据、系统化评估与显式跨视角对齐均对推动MLLM从单视角感知迈向真实世界空间智能至关重要。
原文摘要 · Abstract (English)
Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, visibility, geometry, and interactions across multiple viewpoints. However, progress in cross-view reasoning remains limited by three major gaps: the scarcity of large-scale well-annotated training data, the lack of comprehensive benchmarks for systematic evaluation, and the absence of explicit alignment mechanisms that establish object-level consistency across views. To address these gaps, we thoroughly develop CrossView Suite across three coordinated components: CrossViewSet, CrossViewBench, and CrossViewer. Firstly, we introduce a multi-agent data engine to meticulously curate a large-scale, high-quality cross-view instruction dataset, termed CrossViewSet, covering 17 fine-grained task types with 1.6M samples. Second, we meticulously create a scene-disjoint CrossViewBench to comprehensively assess the cross-view spatial understanding capability of an MLLM, evaluating it across various aspects. Finally, we propose CrossViewer, a progressive three-stage framework for cross-view spatial reasoning in MLLMs, following a Perception -> Alignment -> Reasoning paradigm. Our method equips an adaptive spatial region tokenizer to capture fine-grained object representations, and then aligns the multi-view objects explicitly, and thus fuses aligned features for boosting the cross-view inference capacity for MLLMs. Extensive experiments and analyses show that large-scale training data, systematic evaluation, and explicit cross-view alignment are all critical for advancing MLLMs from single-view perception toward real-world spatial intelligence. The project page is available at https://github.com/Thinkirin/Crossview-Suite.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。