arXiv:2512.04563cs.CV2025-12被引 10

统一模型提升空间感知与推理,让大模型更懂物体位置关系。

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

论文配图:COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
图 1 · 摘自论文原文
  • 用深度和分割图作辅助信息,训练统一模型同时提升感知与推理。
  • 空间推理平均提升6.91%,距离大小估计提升7.92%。
  • 适合做多模态理解、自动驾驶等需要空间智能的场景。

视觉空间推理对多模态大语言模型理解物体属性与空间关系至关重要,但现有模型仍缺乏3D感知能力。当前方法通常分别增强感知(如添加深度、分割图)或推理(如在空间VQA上训练并使用强化学习),将二者割裂。本文提出统一模型COOPER,利用深度与分割作为辅助模态,并分两阶段训练:第一阶段学习生成辅助模态,第二阶段实现自适应交错推理。实验显示,COOPER在空间推理上平均提升6.91%,且仅训练生成辅助模态的变体在距离与尺寸估计上仍达7.92%提升,表明生成辅助模态有助于内化空间知识,增强空间理解能力。

原文摘要 · Abstract (English)

Visual Spatial Reasoning is crucial for enabling Multimodal Large Language Models (MLLMs) to understand object properties and spatial relationships, yet current models still struggle with 3D-aware reasoning. Existing approaches typically enhance either perception, by augmenting RGB inputs with auxiliary modalities such as depth and segmentation, or reasoning, by training on spatial VQA datasets and applying reinforcement learning, and thus treat these two aspects in isolation. In this work, we investigate whether a unified MLLM can develop an intrinsic ability to enhance spatial perception and, through adaptive interleaved reasoning, achieve stronger spatial intelligence. We propose \textbf{COOPER}, a unified MLLM that leverages depth and segmentation as auxiliary modalities and is trained in two stages to acquire auxiliary modality generation and adaptive, interleaved reasoning capabilities. COOPER achieves an average \textbf{6.91\%} improvement in spatial reasoning while maintaining general performance. Moreover, even a variant trained only for auxiliary modality generation attains a \textbf{7.92\%} gain on distance and size estimation, suggesting that learning to generate auxiliary modalities helps internalize spatial knowledge and strengthen spatial understanding.

空间推理多模态大模型感知融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。