arXiv:2604.20806cs.CVcs.AI2026-04ACL被引 1

评测大模型在多图奥数题中的推理能力,发现最强模型仅达50%准确率

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model

论文配图:OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
图 1 · 摘自论文原文
  • 设计多图像奥数推理基准,证据分散于多图中
  • 顶尖模型如Gemini-3-Pro在该基准上准确率仅约50%
  • 适合研究多图上下文理解与模型改进的团队使用

大型视觉语言模型(LVLM)在奥数级推理任务上已取得显著进展。然而,当前针对这类模型的奥数级多模态推理基准多侧重单图分析,未能充分挖掘多图间的上下文信息。我们提出OMIBench,一个评估奥数级推理能力的基准,当所需证据分布在多张图片时尤为关键。该基准涵盖生物学、化学、数学和物理学奥赛题目,附带人工标注的推理过程及精确匹配与语义匹配的评估协议。在大量实验中,我们观察到现有模型存在明显性能差距。即使是最强的LVLM如Gemini-3-Pro,在该基准上的表现也仅约为50%。这些结果表明,OMIBench可作为研究和提升LVLM多图像推理能力的聚焦资源。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have made substantial advances in reasoning tasks at the Olympiad level. Nevertheless, current Olympiad-level multimodal reasoning benchmarks for these models often emphasize single-image analysis and fail to exploit contextual information across multiple images. We present OMIBench, a benchmark designed to evaluate Olympiad-level reasoning when the required evidence is distributed over multiple images. It contains problems from biology, chemistry, mathematics, and physics Olympiads, together with manually annotated rationales and evaluation protocols for both exact and semantic answer matching. Across extensive experiments on OMIBench, we observe meaningful performance gaps in existing models. Even the strongest LVLMs, such as Gemini-3-Pro, attain only about 50% on the benchmark. These results position OMIBench as a focused resources for studying and improving multi-image reasoning in LVLMs.

多图推理奥数题视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。