评测40个多模态模型在化学奥赛题上的表现,发现视觉语言融合常出错。
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- 用化学奥赛题测试多模态模型的图文推理能力
- 部分模型删图后准确率反而提升,说明视觉融合有误
- 思维链提示能显著提高准确率和可解释性
多模态科学推理仍是大语言模型的重大挑战,尤其在化学领域,问题解决依赖符号图、分子结构和结构化视觉数据。本文系统评估了40个专有及开源多模态大模型(包括GPT-5、o3、Gemini-2.5-Pro、Qwen2.5-VL),在涵盖二十年美国国家化学奥林匹克竞赛(USNCO)题目的定制化基准上进行测试。这些题目要求跨多种模态的综合图文推理。结果显示,许多模型在模态融合上表现不佳,某些情况下移除图像反而提升准确率,表明视觉-语言对齐存在缺陷。通过消融实验和遮蔽可解释性分析,证明思维链(Chain-of-Thought)提示能持续提升准确率与视觉定位能力。研究揭示了当前多模态大模型在科学推理方面的关键局限,为开发更鲁棒、可解释的化学多模态系统提供了可行策略。本工作为领域特定多模态人工智能的进步提供及时基准,并强调了人工智能与科学推理交叉领域的进一步发展需求。
原文摘要 · Abstract (English)
Multimodal scientific reasoning remains a significant challenge for large language models (LLMs), particularly in chemistry, where problem-solving relies on symbolic diagrams, molecular structures, and structured visual data. Here, we systematically evaluate 40 proprietary and open-source multimodal LLMs, including GPT-5, o3, Gemini-2.5-Pro, and Qwen2.5-VL, on a curated benchmark of Olympiad-style chemistry questions drawn from over two decades of U.S. National Chemistry Olympiad (USNCO) exams. These questions require integrated visual and textual reasoning across diverse modalities. We find that many models struggle with modality fusion, where in some cases, removing the image even improves accuracy, indicating misalignment in vision-language integration. Chain-of-Thought prompting consistently enhances both accuracy and visual grounding, as demonstrated through ablation studies and occlusion-based interpretability. Our results reveal critical limitations in the scientific reasoning abilities of current MLLMs, providing actionable strategies for developing more robust and interpretable multimodal systems in chemistry. This work provides a timely benchmark for measuring progress in domain-specific multimodal AI and underscores the need for further advances at the intersection of artificial intelligence and scientific reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。