动态科学理解评测基准,实时追踪前沿科研进展
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- 构建可随科研发展更新的活体评测集,覆盖2.5万组图文对
- 现有模型在跨模态科学推理上表现有限,最高提升11%
- 适合关注科学智能与模型评估的科研人员使用
随着多模态大语言模型能力不断提升,固定评测基准逐渐难以有效评估其高水平科学理解能力。本文提出多模态学术封面基准(MAC),一个可随科学进步与模型演进持续更新的活体评测集。MAC基于《自然》《科学》《细胞》等顶级期刊的近2.5万组图像-文本对,挑战模型在抽象视觉与文本内容间的跨模态科学推理能力。对最新年度快照MAC-2025的实验显示,尽管模型具备较强感知能力,其跨模态科学推理仍受限。为此,我们提出DAD——一种轻量级推理时增强方法,通过将视觉特征扩展至语言空间进行推理,实现最高达11%的性能提升。最后,通过更新期刊封面与模型筛选机制的实验证明了MAC的动态特性,展示了其持续与人类知识前沿对齐的潜力。基准代码已开源。
原文摘要 · Abstract (English)
As multimodal large language models (MLLMs) grow increasingly capable, fixed benchmarks are gradually losing their effectiveness in evaluating high-level scientific understanding. In this paper, we introduce the Multimodal Academic Cover benchmark (MAC), a live benchmark that could continuously evolve with scientific advancement and model progress. MAC leverages over 25,000 image-text pairs sourced from issues of top-tier scientific journals such as Nature, Science, and Cell, challenging MLLMs to reason across abstract visual and textual scientific content. Experiments on our most recent yearly snapshot, MAC-2025, reveal that while MLLMs demonstrate strong perceptual abilities, their cross-modal scientific reasoning remains limited. To bridge this gap, we propose DAD, a lightweight inference-time approach that enhances MLLMs by extending MLLM visual features with language space reasoning, achieving performance improvements of up to 11%. Finally, we highlight the live nature of MAC through experiments on updating journal covers and models for curation, illustrating its potential to remain aligned with the frontier of human knowledge. We release our benchmark at https://github.com/mhjiang0408/MAC_Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。