arXiv:2503.01891cs.LGcs.CL2025-03ACL被引 9

评测中文多模态科学问题的推理能力,发现顶尖模型准确率仅63.77%。

MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems

  • 构建包含文本与图文格式的中文多模态科学题集,标注难度与解题过程。
  • 最先进模型在图文任务中准确率仅为63.77%,视觉推理能力薄弱。
  • 适合关注多模态科学推理、模型评估的研究者使用。

大型语言模型(LLMs)和视觉-语言模型(LVLMs)虽在多项任务中表现优异,但其科学推理能力尚未充分检验,尤其在多模态场景下。本文提出MMSciBench,一个用于评估数学与物理推理能力的基准,涵盖纯文本与图文双格式,包含人工标注的难度等级、带详细解释的解题过程及分类映射。对现有顶尖模型的评估显示存在显著局限:即使最佳模型准确率也仅为63.77%,且在视觉推理任务中表现尤为不佳。分析揭示了复杂推理与图文融合方面的关键差距,确立了MMSciBench作为衡量多模态科学理解进展的严格标准。代码已开源至GitHub,数据集可在Hugging Face获取。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) and vision-language models (LVLMs) have shown promise across many tasks, yet their scientific reasoning capabilities remain untested, particularly in multimodal settings. We present MMSciBench, a benchmark for evaluating mathematical and physical reasoning through text-only and text-image formats, with human-annotated difficulty levels, solutions with detailed explanations, and taxonomic mappings. Evaluation of state-of-the-art models reveals significant limitations, with even the best model achieving only \textbf{63.77\%} accuracy and particularly struggling with visual reasoning tasks. Our analysis exposes critical gaps in complex reasoning and visual-textual integration, establishing MMSciBench as a rigorous standard for measuring progress in multimodal scientific understanding. The code for MMSciBench is open-sourced at GitHub, and the dataset is available at Hugging Face.

多模态科学推理评测基准中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。