首个专测视觉遮挡理解的基准,发现大模型仍远差于人类。
Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- 用分层合成生成1365张遮挡图像,构建多任务问答数据集
- 22个模型测试显示与人类差距显著,量化任务尤其弱
- 揭示保守倾向、整体感知脆弱等三大失败模式
遮挡感知是实现人类级空间理解的关键,涉及视觉识别与推理的融合。尽管多模态大语言模型(MLLMs)表现卓越,其在遮挡感知上的能力仍缺乏系统评估。为此,我们提出O-Bench,首个专注于遮挡感知的视觉问答(VQA)基准。基于SA-1B数据集,采用新颖的分层合成方法构建1,365张语义一致的遮挡场景图像,并通过可靠半自动流程标注4,588组问题答案,涵盖五个定制任务。对22个代表性MLLM的全面评估表明,当前模型与人类基准存在显著性能差距,且该差距无法通过模型规模扩大或思维链增强有效缓解。我们进一步识别出三种典型失败模式:过度保守倾向、脆弱的整体知觉预测以及在定量任务中的困难。我们认为O-Bench不仅能为遮挡感知提供关键评估工具,还将推动MLLM视觉智能的发展。该基准将在论文发表后公开。
原文摘要 · Abstract (English)
Occlusion perception, a critical foundation for human-level spatial understanding, embodies the challenge of integrating visual recognition and reasoning. Though multimodal large language models (MLLMs) have demonstrated remarkable capabilities, their performance on occlusion perception remains under-explored. To address this gap, we introduce O-Bench, the first visual question answering (VQA) benchmark specifically designed for occlusion perception. Based on SA-1B, we construct 1,365 images featuring semantically coherent occlusion scenarios through a novel layered synthesis approach. Upon this foundation, we annotate 4,588 question-answer pairs in total across five tailored tasks, employing a reliable, semi-automatic workflow. Our extensive evaluation of 22 representative MLLMs against the human baseline reveals a significant performance gap between current MLLMs and humans, which, we find, cannot be sufficiently bridged by model scaling or thinking process. We further identify three typical failure patterns, including an overly conservative bias, a fragile gestalt prediction, and a struggle with quantitative tasks. We believe O-Bench can not only provide a vital evaluation tool for occlusion perception, but also inspire the development of MLLMs for better visual intelligence. Our benchmark will be made publicly available upon paper publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。