arXiv:2411.16955cs.LGcond-mat.mtrl-sci2024-11被引 73

测试多模态模型在化学材料研究中的局限性,发现其推理能力不足。

Probing the limitations of multimodal language models for chemistry and materials research

  • 构建了针对化学材料任务的多模态评测基准MaCBench
  • 模型在设备识别上接近完美,但空间推理能力差
  • 适合关注多模态AI在科研中应用瓶颈的研究者

人工智能的进展激发了科学助手的兴趣,这类系统可支持从文献综述到实验设计与数据分析的全流程工作。关键能力在于处理和推理文本与视觉形式的科学信息——从解析光谱数据到理解实验装置。本文提出MaCBench,一个全面的基准,用于评估视觉-语言模型在真实化学与材料科学任务中的表现,涵盖数据提取、实验理解与结果解释三个核心方面。对领先模型的系统评估显示,尽管在基础感知任务中表现良好(如设备识别几乎完美,标准数据提取准确率高),但在空间推理、跨模态信息融合及多步逻辑推理方面存在根本性局限。这些发现对化学与材料科学之外也具启示意义,表明开发可靠的多模态科学助手需在训练数据的收集与训练方法上取得突破。

原文摘要 · Abstract (English)

Recent advancements in artificial intelligence have sparked interest in scientific assistants that could support researchers across the full spectrum of scientific workflows, from literature review to experimental design and data analysis. A key capability for such systems is the ability to process and reason about scientific information in both visual and textual forms - from interpreting spectroscopic data to understanding laboratory setups. Here, we introduce MaCBench, a comprehensive benchmark for evaluating how vision-language models handle real-world chemistry and materials science tasks across three core aspects: data extraction, experimental understanding, and results interpretation. Through a systematic evaluation of leading models, we find that while these systems show promising capabilities in basic perception tasks - achieving near-perfect performance in equipment identification and standardized data extraction - they exhibit fundamental limitations in spatial reasoning, cross-modal information synthesis, and multi-step logical inference. Our insights have important implications beyond chemistry and materials science, suggesting that developing reliable multimodal AI scientific assistants may require advances in curating suitable training data and approaches to training those models.

多模态科学智能化学AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。