构建多语言科学问答基准,挑战大模型跨模态推理能力
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- 设计三类评估模式,覆盖四大学科五种语言
- 图像单模态下最高准确率仅52.11%,显著高于现有基准
- 细粒度标注揭示模型在数学物理等领域的薄弱环节
近年来,多模态大语言模型(MLLMs)在多个领域取得显著进展,相应评估基准也持续优化。然而,现有科学领域基准仍面临三大挑战:1)多语言场景下模型推理能力评估不足;2)多模态覆盖不全面;3)科学知识点标注粗略。为此,我们提出MME-SCI,一个全面且具有挑战性的科学评测基准。共收集1,019个高质量问答对,涵盖数学、物理、化学、生物四大学科,支持中文、英文、法文、西班牙文、日文五种语言,并包含三种评估模式。在16个开源与4个闭源模型上进行广泛实验,结果显示现有模型在该基准上表现普遍不佳。例如,在仅图像输入的评估模式下,o4-mini在数学、物理、化学、生物学上的准确率分别为52.11%、24.73%、36.57%和29.80%,显著高于以往基准。更重要的是,利用其多语言与细粒度知识标注特性,深入分析模型表现,识别出其在特定学科中的弱点。数据与评估代码已开源。
原文摘要 · Abstract (English)
Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the reasoning capabilities of MLLMs. However, existing benchmarks still face three key challenges: 1) Insufficient evaluation of models' reasoning abilities in multilingual scenarios; 2) Inadequate assessment of MLLMs' comprehensive modality coverage; 3) Lack of fine-grained annotation of scientific knowledge points. To address these gaps, we propose MME-SCI, a comprehensive and challenging benchmark. We carefully collected 1,019 high-quality question-answer pairs, which involve 3 distinct evaluation modes. These pairs cover four subjects, namely mathematics, physics, chemistry, and biology, and support five languages: Chinese, English, French, Spanish, and Japanese. We conducted extensive experiments on 16 open-source models and 4 closed-source models, and the results demonstrate that MME-SCI is widely challenging for existing MLLMs. For instance, under the Image-only evaluation mode, o4-mini achieved accuracy of only 52.11%, 24.73%, 36.57%, and 29.80% in mathematics, physics, chemistry, and biology, respectively, indicating a significantly higher difficulty level compared to existing benchmarks. More importantly, using MME-SCI's multilingual and fine-grained knowledge attributes, we analyzed existing models' performance in depth and identified their weaknesses in specific domains. The Data and Evaluation Code are available at https://github.com/JCruan519/MME-SCI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。