首个全景空间推理评测基准,揭示大模型在360度场景下推理能力不足。
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
- 构建首个面向全景图像的空间推理评测集,含超15万题对
- 八款主流大模型在零样本下表现均不佳,空间推理准确率低
- 提出抗幻觉采样策略和双阶段评估框架,适合研究视觉-语言模型的学者
360度摄像头捕捉的180x360全景视场广泛应用于具身智能与虚拟现实。尽管多模态大语言模型在标准针孔图像的视觉-空间推理上取得进展,但对全景感知的研究仍不足。本文提出核心问题:多模态大模型是否已准备好处理全景空间推理?为此,我们构建了首个专门针对该场景的评测基准OSR-Bench,包含超过15.3万组基于高保真全景室内场景地图的问题-答案对,覆盖物体计数、相对距离和方向等关键推理类型。我们设计了一种负采样策略,在提示中插入不存在物体以评估幻觉与定位鲁棒性。为实现细粒度分析,提出两阶段评估框架,结合旋转不变匹配与规则+LLM混合指标,评估认知地图生成与问答准确率。在零样本设置下,对包括GPT-4o、Gemini 1.5 Pro在内的八款先进模型进行测试,结果显示当前模型在全景上下文中空间推理能力普遍薄弱,凸显亟需更贴近感知的多模态模型。OSR-Bench及代码将公开于:https://huggingface.co/datasets/UUUserna/OSR-Bench
原文摘要 · Abstract (English)
The 180x360 omnidirectional field of view captured by 360-degree cameras enables their use in a wide range of applications such as embodied AI and virtual reality. Although recent advances in multimodal large language models (MLLMs) have shown promise in visual-spatial reasoning, most studies focus on standard pinhole-view images, leaving omnidirectional perception largely unexplored. In this paper, we ask: Are MLLMs ready for omnidirectional spatial reasoning? To investigate this, we introduce OSR-Bench, the first benchmark specifically designed for this setting. OSR-Bench includes over 153,000 diverse question-answer pairs grounded in high-fidelity panoramic indoor scene maps. It covers key reasoning types including object counting, relative distance, and direction. We also propose a negative sampling strategy that inserts non-existent objects into prompts to evaluate hallucination and grounding robustness. For fine-grained analysis, we design a two-stage evaluation framework assessing both cognitive map generation and QA accuracy using rotation-invariant matching and a combination of rule-based and LLM-based metrics. We evaluate eight state-of-the-art MLLMs, including GPT-4o, Gemini 1.5 Pro, and leading open-source models under zero-shot settings. Results show that current models struggle with spatial reasoning in panoramic contexts, highlighting the need for more perceptually grounded MLLMs. OSR-Bench and code will be released at: https://huggingface.co/datasets/UUUserna/OSR-Bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。