多机器人协作空间推理新框架,让AI理解多个视角的动态环境。
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

- 用多视角视频融合与物理约束引导推理,提升协同感知能力。
- 在Habitat和iGibson上分别提升3.87%和7.12%准确率。
- 适用于真实四足机器人测试,可推广到未知团队规模。
多模态大语言模型在第一人称视频理解上已取得显著进展,但其从多个具身视角协同推理的能力仍待探索。本文通过多机器人协作动态空间推理任务研究该问题,要求模型结合同步的第一人称视频流回答空间、时间、可见性及协调性问题。为此,我们提出CoopSR基准与EgoTeam数据集,包含114,227个问答对,涵盖19种题型、4个难度层级和3种团队规模,在Habitat和iGibson仿真环境及约2,326个真实四足机器人采集的测试集上构建。我们进一步提出SP-CoR(谱与物理引导的协作推理器)框架,结合动态感知的多机器人帧采样、谱与物理引导的视图融合及物理对齐的提示蒸馏,使模型在训练中利用机器人位姿监督,测试时仅需第一人称视频。在22个基线模型上,SP-CoR持续提升协作推理性能,在Habitat和iGibson上分别超越最强微调基线+3.87%和+7.12%,并展现出对未见团队规模和真实场景更强的泛化能力。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study this problem through multi-robot cooperative dynamic spatial reasoning, where a model must answer spatial, temporal, visibility, and coordination questions by integrating synchronized egocentric videos from a team of moving robots. To support this setting, we introduce CoopSR, the first benchmark for this task, together with EgoTeam, a multi-robot egocentric QA dataset. EgoTeam contains 114,227 QA pairs spanning 19 question types, four difficulty tiers, and three team sizes in Habitat and iGibson, along with a real-world test set of around 2,326 QAs collected using two quadruped robots. We further propose SP-CoR (Spectral and Physics-Informed Cooperative Reasoner), an MLLM framework for fine-grained cooperative spatial reasoning. SP-CoR combines dynamics-aware multi-robot frame sampling, spectral- and physics-guided view fusion, and physics-aligned prompt distillation, enabling the model to benefit from privileged robot-pose supervision during training while requiring only egocentric videos at test time. Across 22 MLLM baselines, SP-CoR consistently improves cooperative reasoning, outperforming the strongest fine-tuned baseline by +3.87% on Habitat and +7.12% on iGibson. It also shows stronger generalization to unseen team sizes and real-world robot tests. Code can be found at https://github.com/KPeng9510/seeing-together.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。