arXiv:2509.15839cs.CL2025-09

构建中文物理多模态推理基准,评估模型看图解题能力

Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems

  • 设计5级难度、覆盖11个高中物理主题的多模态题目集
  • 20个模型在图像关联题上平均准确率仅43.7%,显示严重短板
  • 适合研究多模态推理、中文科学教育的学者与开发者

尽管多模态大模型(MLLMs)在推理任务中取得显著进展,但在物理等专业科学领域仍存在评估基准不足的问题。现有基准普遍存在主题覆盖不全、忽略逐步推理过程、以英文为主等问题,难以系统评估视觉信息的作用。为此,我们提出面向中文物理推理的综合性基准Multi-Physics,包含5个难度等级,涵盖11个高中物理主题,共1,412道图文结合的单选题。我们采用双维度评估框架,对20种不同MLLM进行测试,分析其最终答案准确率与思维链(chain-of-thought)完整性。此外,通过对比输入模式改变前后的表现,系统研究难度层级与视觉信息的影响。本工作不仅为社区提供细粒度资源,还建立了一套剖析前沿多模态模型推理过程的可靠方法。数据集与代码已开源:https://github.com/luozhongze/Multi-Physics。

原文摘要 · Abstract (English)

While multimodal LLMs (MLLMs) demonstrate remarkable reasoning progress, their application in specialized scientific domains like physics reveals significant gaps in current evaluation benchmarks. Specifically, existing benchmarks often lack fine-grained subject coverage, neglect the step-by-step reasoning process, and are predominantly English-centric, failing to systematically evaluate the role of visual information. Therefore, we introduce \textbf {Multi-Physics} for Chinese physics reasoning, a comprehensive benchmark that includes 5 difficulty levels, featuring 1,412 image-associated, multiple-choice questions spanning 11 high-school physics subjects. We employ a dual evaluation framework to evaluate 20 different MLLMs, analyzing both final answer accuracy and the step-by-step integrity of their chain-of-thought. Furthermore, we systematically study the impact of difficulty level and visual information by comparing the model performance before and after changing the input mode. Our work provides not only a fine-grained resource for the community but also offers a robust methodology for dissecting the multimodal reasoning process of state-of-the-art MLLMs, and our dataset and code have been open-sourced: https://github.com/luozhongze/Multi-Physics.

多模态推理中文评测物理教育链式思考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。