构建更严格的多模态理解测试,逼模型真正‘看懂图文’
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
- 三步优化:筛除纯文本可解题、扩增选项、图文混合输入
- 模型在新基准上性能下降16.8%至26.9%,体现真实能力差距
- 适合关注多模态推理与视觉语言融合研究的开发者
本文提出MMMUS-Pro,是大规模多学科多模态理解与推理基准MMMU的强化版本。通过三阶段严格评估:(1) 剔除仅靠文本即可回答的问题,(2) 扩充候选答案,(3) 引入视觉输入仅限图像的设置,将问题嵌入图像中。该设定要求模型同时“看见”和“读取”信息,检验人类核心认知能力——跨模态融合。结果显示,各模型在MMMU-Pro上的表现显著低于原始基准,得分范围为16.8%至26.9%。我们研究了OCR提示与思维链(CoT)推理的影响,发现OCR提示效果有限,而CoT普遍提升性能。MMMU-Pro提供更贴近真实场景的评估方式,为未来多模态人工智能研究指明方向。
原文摘要 · Abstract (English)
This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly "see" and "read" simultaneously, testing a fundamental human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future research in multimodal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。