arXiv:2511.13719cs.CVcs.AI2025-11中稿 · CVPR被引 71

通过800万空间智能数据训练,提升多模态大模型的空间理解能力。

Scaling Spatial Intelligence with Multimodal Foundation Models

论文配图:Scaling Spatial Intelligence with Multimodal Foundation Models
图 1 · 摘自论文原文
  • 构建800万条空间能力数据集,系统化提升模型空间推理能力
  • 在多个基准上达到68.8%至85.7%的领先性能,同时保持强多模态理解
  • 适合研究空间智能、多模态模型训练及下游应用的开发者与研究者

尽管取得显著进展,多模态基础模型在空间智能方面仍存在明显缺陷。本文通过系统性构建涵盖八百万多样化数据样本的SenseNova-SI-8M数据集,基于现有视觉理解模型(如Qwen3-VL、InternVL3)和统一理解生成模型(如Bagel),在SenseNova-SI系列中培养空间智能。该模型在多项空间智能基准上表现卓越:VSI-Bench达68.8%,MMSI达43.3%,MindCube达85.7%,ViewSpatial达54.7%,SITE达47.7%,BLINK达63.9%,3DSR达55.5%,EmbSpatial达72.0%;同时保持强大通用多模态理解能力(MMBench-En达84.9%)。我们分析了数据规模的影响,发现多样数据训练可催生涌现式泛化能力,揭示过拟合与语言捷径风险,并初步验证空间链式思维推理潜力。所有新模型均已公开发布。

原文摘要 · Abstract (English)

Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation models to cultivate spatial intelligence within the SenseNova-SI family, built upon established multimodal foundations including visual understanding models (i.e., Qwen3-VL and InternVL3) and unified understanding and generation models (i.e., Bagel). We take a principled approach to constructing high-performing and robust spatial intelligence by systematically curating SenseNova-SI-8M: eight million diverse data samples under a rigorous taxonomy of spatial capabilities. SenseNova-SI demonstrates unprecedented performance across a broad range of spatial intelligence benchmarks: 68.8% on VSI-Bench, 43.3% on MMSI, 85.7% on MindCube, 54.7% on ViewSpatial, 47.7% on SITE, 63.9% on BLINK, 55.5% on 3DSR, and 72.0% on EmbSpatial, while maintaining strong general multimodal understanding (e.g., 84.9% on MMBench-En). More importantly, we analyze the impact of data scaling, discuss early signs of emergent generalization capabilities enabled by diverse data training, analyze the risk of overfitting and language shortcuts, present a preliminary study on spatial chain-of-thought reasoning, and validate the potential downstream application. All newly trained multimodal foundation models are publicly released.

空间智能多模态模型数据训练基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。