arXiv:2605.29833cs.AI2026-05被引 1

构建跨19个材料科学领域的多模态推理基准,评估AI在科研中的实际推理能力。

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

论文配图:OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
图 1 · 摘自论文原文
  • 设计覆盖19个子领域的专家标注问答与计算题集
  • 13个模型最高仅得0.372分,显示现有AI推理能力严重不足
  • 揭示知识分布不均、思维模式固化等关键缺陷,适合研究AI科研助手

随着多模态语言模型在科研中作用日益重要,材料科学因其跨学科性、多模态特征和应用导向,成为理想测试场景。然而现有材料基准多集中于性质预测、知识问答或表征理解,缺乏对从材料知识到应用的完整推理过程的考察。为此,我们提出OmniMatBench,一个经过人工校准的材料科学多模态推理基准。该基准包含3,171道专家标注的问答与计算题,覆盖19个材料科学子领域,涵盖基础材料知识、结构与工程材料、制备与制造工艺、功能与应用材料。我们评估了13个开源与闭源多模态大模型(MLLM),发现表现最佳模型仅得0.372分,暴露出当前材料科学推理能力的巨大差距。进一步分析显示各子领域表现差异显著,存在固定推理策略、材料知识分布不均、高阶知识应用受限等问题,在公式、检索与代码辅助设置下尤为明显。OmniMatBench为理解当前多模态大模型在材料科学中的能力边界提供了关键洞察,并为构建可靠的科研人工智能助手奠定了基础。

原文摘要 · Abstract (English)

As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdisciplinary, multimodal, and application-driven nature. However, existing materials benchmarks mainly focus on property prediction, knowledge QA, or characterization understanding, leaving the broader reasoning process from materials knowledge to application underexplored. To fill this gap, we present OmniMatBench, a human-calibrated multimodal reasoning benchmark for materials science. OmniMatBench contains 3,171 expert-curated QA and calculation problems across 19 materials-science subfields, spanning fundamental materials knowledge, structural and engineering materials, materials processing and manufacturing, and functional and applied materials. We evaluate 13 open-source and closed-source MLLMs and find that the best model achieves only a 0.372 overall score, revealing a substantial gap in current materials-science reasoning. Further analysis shows strong variation across subfields, fixed reasoning heuristics, uneven materials knowledge, and limited high-level knowledge application under formula-, retrieval-, and code-assisted settings. OmniMatBench provides crucial insights into the capabilities and limitations of current MLLMs and establishes a foundation for reliable AI assistants in materials-science research.

多模态推理材料科学基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。