arXiv:2609.07258cs.CVcs.AI2026-09

构建分层多视角空间推理数据集,提升大模型3D空间理解能力

MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling

论文配图:MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling
图 1 · 摘自论文原文
  • 设计分层依赖结构,模拟人类空间认知的渐进推理路径
  • 生成跨视图强约束的多级推理题目,支持链式思维监督
  • 在多视角基准上实现领先效果,适合需要3D理解的任务

尽管多模态大模型在二维视觉-语言任务中进展迅速,但多视角空间推理仍因现有数据集缺乏结构化3D认知路径而面临根本瓶颈。为此,我们提出MV-STRIDE——一个具有相互依赖与分解能力的多视角空间推理数据集。该数据集超越扁平数据结构,显式建模基础感知、场景理解与复杂上下文推理之间的依赖关系,提供与人类空间认知一致的学习路径。我们开发了一套系统化的问答生成流程,利用多样化的3D场景源,强制施加跨视图依赖约束,防止单视图可解,生成由认知基础链式思维监督支持的多层级空间推理任务。大量评估表明,基于该分层数据集的多阶段训练框架在多个空间推理基准上达到当前最优性能,尤其在面向多视角的MMSI-Bench上表现突出。本方法使多模态大模型能够在多样化视角下保持稳健且一致的3D空间推理能力。代码与数据集已公开于https://co1dspring.github.io/MV-STRIDE/。

原文摘要 · Abstract (English)

Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust multi-view spatial reasoning remains a fundamental bottleneck due to the lack of structured 3D cognitive pathways in existing datasets. To address this, we introduce MV-STRIDE, a Multi-View hierarchical SpaTial Reasoning dataset with Interdependent and DEcomposed capabilitiEs. Moving beyond flat data structures, MV-STRIDE explicitly models the dependency relationships between foundational perception, scene understanding, and complex contextual reasoning, providing a coherent learning pathway aligned with human spatial cognition. We develop a systematic QA generation pipeline leveraging diverse 3D scene sources that enforces cross-view dependency constraints to prevent single-view solvability, generating multi-level spatial reasoning tasks supported by cognitively grounded chain-of-thought supervision for complex inference. Extensive evaluations demonstrate that our multi-stage training framework based on our hierarchical dataset achieves state-of-the-art performance across multiple spatial reasoning benchmarks, notably the multi-view oriented MMSI-Bench. Our approach enables MLLMs to maintain robust, 3D-consistent spatial reasoning across diverse viewpoints. The code and dataset are available at https://co1dspring.github.io/MV-STRIDE/.

空间推理多视角大模型3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。