arXiv:2511.09339cs.CL2025-11中稿 · IJCNLP-AACL Findin…被引 2

评测视觉语言模型科学推理能力,用印度高考真题区分真实理解与模式匹配。

mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models

  • 基于印度理工科高考题构建双语多模态评测集,覆盖物理、化学、数学。
  • 顶尖模型在2025年新题上达77-84%准确率,开源模型仅37-45%。
  • 揭示模型在深度推理和元认知压力下的崩溃,适合评估真正科学推理能力。

当前视觉语言模型(VLMs)在现有多模态推理基准上表现良好(如MMMU、MathVista准确率78-85%),但难以区分真正的科学推理能力与模式匹配。为填补这一空白,我们提出**mmJEE-Eval**,一个包含1,460道题的双语(英语与印地语)多模态基准,题目来自印度2019-2025年JEE Advanced考试,涵盖高中阶段的物理、化学与数学领域。对17个前沿模型的评估显示,领先模型(GPT-5、Gemini 2.5 Pro/Flash)在未见的2025年题目上达到77-84%准确率,而开源模型虽扩展至400B参数仍停滞于37-45%。尽管闭源模型在高阶问题上表现优异(最高pass@3达100%),但在元认知推理压力下全面崩溃(如GPT-5仅修正5.2%错误)。系统性消融分析表明,该基准的难度源于推理深度与复杂性,而非记忆。该评测有效区分了先进的训练与推理方法,验证了其筛选能力。代码与数据已公开:https://mmjee-eval.github.io

原文摘要 · Abstract (English)

Contemporary vision-language models (VLMs) perform well on existing multimodal reasoning benchmarks (78-85\% accuracy on MMMU, MathVista). Yet, these results fail to sufficiently distinguish true scientific reasoning articulation capabilities from pattern-matching. To address this gap, we introduce \textbf{mmJEE-Eval}, a multimodal bilingual (English and Hindi) benchmark comprising 1,460 questions from India's JEE Advanced examination (2019-2025) spanning pre-college Physics, Chemistry, and Mathematics domains. Our evaluation of 17 state-of-the-art models reveals that while frontier VLMs (GPT-5, Gemini 2.5 Pro/Flash) achieve 77-84\% accuracy on held-out 2025 questions, open-source models plateau at 37-45\% despite scaling to 400B parameters, a significant difference not observed on existing benchmarks. While closed frontiers from Google and OpenAI show high problem-solving accuracies (up to 100\% pass@3 scores), they fully collapse when the reasoning load is increased meta-cognitively (GPT-5 fixes just 5.2\% errors). Systematic ablations show mmJEE-Eval's difficulty stems from complexity and reasoning depth rather than memorization. Effectively, our benchmark segregates superior training and reasoning methodologies where alternatives fail. We publicly release our code and data: https://mmjee-eval.github.io

多模态评测科学推理双语模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。