构建高质量多模态评估基准,精准衡量模型真实理解能力。
MMGist: A Comprehensive Multimodal Benchmark for 2027

- 三阶段过滤:文本消融、模型饱和度、异常项检测,剔除无效数据。
- 仅用7262项实现98%模型排名保真度,评估量减少69%,区分度提升78%。
- 揭示视觉逻辑弱、专家知识差异等关键短板,适合模型评测与改进参考。
我们系统分析了18个主流视觉语言基准,发现三大问题:1)大量题目不依赖视觉线索,无法有效衡量多模态理解;2)许多题目已接近当前大型视觉语言模型(LVLM)性能饱和,缺乏区分力;3)少量异常题目影响评估可靠性。为此,我们提出MMGist,一个经过筛选的基准,涵盖七个能力维度,共7,262个样本。该基准通过三阶段流程构建:文本消融过滤、跨模型饱和度过滤和异常检测过滤。我们在27个领先LVLM上进行实验,对比原始23,250项数据池。结果表明,MMGist在保持模型排名高保真度(斯皮尔曼相关系数ρ=0.98)的同时,将评估项减少69%,跨模型区分度提升78%。进一步分析显示,当前LVLM在视觉逻辑方面仍存在系统性短板,而专家知识等知识密集型维度仍是区分闭源与开源自模型的关键因素。研究强调,高质量评估应优先关注视觉依赖性、区分力与可靠性,而非单纯扩大规模。
原文摘要 · Abstract (English)
We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effectively measure multimodal understanding; 2) many items are already close to performance saturation for current LVLMs, which limits their discriminative power; 3) a small number of anomalous items affect the reliability of evaluation results. To this end, we propose MMGist, a curated benchmark that covers seven capability dimensions and contains 7,262 items. MMGist is constructed through a three-stage pipeline, which sequentially combines text-ablation filtering, cross-model saturation filtering, and anomaly detection filtering. We conduct extensive experiments on 27 leading LVLMs and compare MMGist with the raw pool of 23,250 items. The results show that MMGist preserves model rankings with high fidelity, with Spearman $ρ= 0.98$, while reducing evaluation items by 69\% and improving cross-model discrimination by 78\%. Further results indicate that Visual Logic remains a systematic weakness of current LVLMs, while knowledge-intensive dimensions such as Expert Knowledge dimensions remain important factors for distinguishing closed-source models from open-source models. These findings suggest that high-quality evaluation should prioritize visual dependency, discriminative power, and reliability, rather than simply pursuing benchmark scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。