arXiv:2507.07984cs.CV2025-07NeurIPS被引 31

评测多模态大模型在线时空理解能力,揭示其在动态探索中的短板。

OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding

  • 构建在线时空推理基准,模拟智能体逐步观察场景的动态过程。
  • 1.4千场景、1万问答对,模型在长时探索中准确率显著下降。
  • 揭示线索推理与长期记忆是提升在线感知的关键挑战,适合具身智能研究者。

多模态大语言模型在视觉与语言融合推理方面取得显著进展。然而,现有评估大多基于离线固定输入,难以反映真实世界中的动态感知挑战。本文提出OST-Bench,一个从主动探索视角评估模型在线时空理解能力的基准。该基准强调增量观测的处理与历史记忆的整合,以支持动态空间推理。基于高效数据采集流程,数据集包含1.4k场景和10k问答对,来自ScanNet、Matterport3D和ARKitScenes。我们评估多个领先多模态大模型,发现其在复杂时空推理任务上表现不佳:随着探索范围扩大和记忆积累,准确率持续下降。进一步分析揭示两类主要错误模式——线索依赖的空间推理需求与长时记忆检索压力分别导致性能衰减,凸显提升在线具身推理的核心瓶颈。代码、数据集与基准已开源,助力后续研究。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have shown remarkable capabilities in integrating vision and language for complex reasoning. While most existing benchmarks evaluate models under offline settings with a fixed set of pre-recorded inputs, we introduce OST-Bench, a benchmark designed to evaluate Online Spatio-Temporal understanding from the perspective of an agent actively exploring a scene. The Online aspect emphasizes the need to process and reason over incrementally acquired observations, while the Spatio-Temporal component requires integrating current visual inputs with historical memory to support dynamic spatial reasoning. OST-Bench better reflects the challenges of real-world embodied perception. Built on an efficient data collection pipeline, OST-Bench consists of 1.4k scenes and 10k question-answer pairs collected from ScanNet, Matterport3D, and ARKitScenes. We evaluate several leading MLLMs on OST-Bench and observe that they fall short on tasks requiring complex spatio-temporal reasoning. Under the online setting, their accuracy declines as the exploration horizon extends and the memory grows. Through further experimental analysis, we identify common error patterns across models and find that both complex clue-based spatial reasoning demands and long-term memory retrieval requirements significantly drop model performance along two separate axes, highlighting the core challenges that must be addressed to improve online embodied reasoning. To foster further research and development in the field, our codes, dataset, and benchmark are available. Our project page is: https://rbler1234.github.io/OSTBench.github.io/

多模态模型具身智能时空推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。