arXiv:2601.16538cs.CV2026-01被引 2

让大模型持续理解动态3D环境,支持真实机器人部署。

OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding

  • 用有限空间记忆持续更新环境认知,不随输入膨胀。
  • 融合3D点云与语义信息,提升物体定位识别精度。
  • 适合需要长期感知的机器人、自动驾驶等场景。

近年来,研究者越来越关注如何使多模态大语言模型(MLLM)具备空间理解与推理能力。然而,现有方法大多忽视了在不断变化的世界中持续工作的重要性,也缺乏在真实世界具身系统中的部署可能性。本文提出OnlineSI框架,能够基于视频流持续改进对周围环境的空间理解。核心思想是维护一个固定大小的空间记忆以保留过往观测,确保空间记忆规模不随输入累积而增长。我们进一步将3D点云信息与语义信息融合,帮助MLLM更准确地定位和识别场景中的物体。为评估方法,我们引入模糊F₁分数以缓解歧义问题,并在两个代表性数据集上进行测试。实验表明该方法有效,为真实世界具身系统的实现铺平了道路。

原文摘要 · Abstract (English)

In recent years, researchers have increasingly been interested in how to enable Multimodal Large Language Models (MLLM) to possess spatial understanding and reasoning capabilities. However, most existing methods overlook the importance of the ability to continuously work in an ever-changing world, and lack the possibility of deployment on embodied systems in real-world environments. In this work, we introduce OnlineSI, a framework that can continuously improve its spatial understanding of its surroundings given a video stream. Our core idea is to maintain a finite spatial memory to retain past observations, ensuring the size of the spatial memory does not increase as the input accumulates. We further integrate 3D point cloud information with semantic information, helping MLLM to better locate and identify objects in the scene. To evaluate our method, we introduce the Fuzzy $F_1$-Score to mitigate ambiguity, and test our method on two representative datasets. Experiments demonstrate the effectiveness of our method, paving the way towards real-world embodied systems.

3D理解具身智能大模型在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。