arXiv:2510.11549cs.CV2025-10被引 11

评测大模型对全景图像的理解能力,发现其普遍表现不佳。

ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?

  • 构建包含2000张全景图和4000个问答对的全新评测基准
  • 20个主流大模型在全景理解任务中表现均不理想,平均准确率不足50%
  • 提出无需训练的思维链方法,显著提升模型对沉浸式环境的理解

全景图像(ODIs)提供360×180度的完整视图,广泛应用于虚拟现实、增强现实及具身智能领域。尽管多模态大语言模型(MLLMs)在传统二维图像与视频理解任务中表现出色,但其对全景图像所呈现的沉浸式环境的理解能力仍缺乏系统评估。为此,我们提出ODI-Bench,一个专门针对全景图像理解的综合性新基准。该基准包含2000张高质量全景图像和超过4000个手工标注的问答对,涵盖10个细粒度任务,覆盖通用级与空间级理解。我们对20个代表性MLLMs(包括开源与闭源模型)在封闭式与开放式设置下进行了全面评测。实验结果表明,当前MLLMs在捕捉全景图像提供的沉浸式上下文方面仍存在显著困难。为此,我们进一步提出Omni-CoT——一种无需训练的方法,通过跨文本信息与视觉线索的链式推理,显著提升模型在全景环境中的理解能力。相关基准与代码将公开于https://github.com/ylylyl-sjtu/ODI-Bench。

原文摘要 · Abstract (English)

Omnidirectional images (ODIs) provide full 360x180 view which are widely adopted in VR, AR and embodied intelligence applications. While multi-modal large language models (MLLMs) have demonstrated remarkable performance on conventional 2D image and video understanding benchmarks, their ability to comprehend the immersive environments captured by ODIs remains largely unexplored. To address this gap, we first present ODI-Bench, a novel comprehensive benchmark specifically designed for omnidirectional image understanding. ODI-Bench contains 2,000 high-quality omnidirectional images and over 4,000 manually annotated question-answering (QA) pairs across 10 fine-grained tasks, covering both general-level and spatial-level ODI understanding. Extensive experiments are conducted to benchmark 20 representative MLLMs, including proprietary and open-source models, under both close-ended and open-ended settings. Experimental results reveal that current MLLMs still struggle to capture the immersive context provided by ODIs. To this end, we further introduce Omni-CoT, a training-free method which significantly enhances MLLMs' comprehension ability in the omnidirectional environment through chain-of-thought reasoning across both textual information and visual cues. Both the benchmark and the code will be released at https://github.com/ylylyl-sjtu/ODI-Bench.

全景图像多模态大模型评测视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。