让大模型学会看连续画面的空间关系,提升机器人感知能力。
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- 构建多帧空间理解框架,融合深度、对应和动态感知
- 创建2700万样本的MultiSPA数据集,支持多帧训练
- 可作机器人多帧奖励标注器,适合具身智能研究者
多模态大语言模型在视觉任务中进展迅速,但其空间理解仍局限于单图,难以满足需要多帧推理的现实应用。本文提出一种新框架,通过整合深度感知、视觉对应和动态感知等基础空间能力,赋予MLLM多帧空间理解能力。设计全新数据流水线,构建包含2700万样本的MultiSPA数据集,覆盖多样三维与四维场景,支撑模型训练。同时引入统一评估基准,涵盖广泛空间任务。所提模型Multi-SpatialMLLM在多个基线和商用系统上表现显著领先,展现可扩展、泛化性强的多帧感知能力。进一步观察到多任务协同效应及复杂场景下的涌现空间能力,并验证其作为机器人多帧奖励标注器的应用潜力。
原文摘要 · Abstract (English)
Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for physical-world applications that require multi-frame reasoning. In this paper, we propose a framework to equip MLLMs with multi-frame spatial understanding by integrating fundamental spatial skills, including depth perception, visual correspondence, and dynamic perception. We design a novel data pipeline and collect the MultiSPA dataset of more than 27 million samples spanning diverse 3D and 4D scenes to enable training. Alongside MultiSPA, we introduce a comprehensive benchmark that tests a wide spectrum of spatial tasks under uniform metrics. Our resulting model, Multi-SpatialMLLM, achieves significant gains over baselines and proprietary systems, demonstrating scalable and generalizable multi-frame perception. We further observe multi-task benefits and emergent spatial capabilities in challenging scenarios, and showcase how our model can serve as a multi-frame reward annotator for robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。