为大模型驱动的材料合成智能体提供评估框架,聚焦与实验设备联动场景
Evaluating LLM-based AI agents integrated with materials synthesis tools: the case of atomic layer deposition

- 构建融合实验工具交互的多层级评估体系
- 以原子层沉积为案例验证评估方法可迁移性
- 适合材料智能研发与自动化实验系统开发者
本文综述了基于大语言模型(LLMs)的AI智能体在材料合成中的性能评估策略。在简要介绍当前主流LLM智能体核心技术后,重点总结了材料科学领域特别是材料合成场景下的评估方法,尤其关注模型与实验设备直接集成的应用情境。评估策略涵盖知识与推理基准、工具使用基准,以及涉及与真实实验系统或虚拟工具交互的闭环基准。以原子层沉积(ALD)为例,阐明现有方法既借鉴了材料科学以外的通用评估范式,又具备向其他合成技术推广的潜力。最后,提出一个实用的评估框架,用于在材料合成背景下系统评价LLM的表现。
原文摘要 · Abstract (English)
This work provides an overview of the different strategies that can be used to evaluate the performance of AI models and agents based on large language models (LLMs) for materials synthesis. After providing a brief overview of the key technologies behind the current generation of AI agents based on LLMs, we summarize the different approaches to evaluating these models in the context of materials science and in particular on materials synthesis, with a specific emphasis on scenarios in which the models are directly integrated with experimental tools. We discuss evaluation strategies spanning knowledge and reasoning benchmarks, tool-use benchmarks, and closed loop benchmarks involving the interaction with experimental systems or realistic virtual tools. We use atomic layer deposition (ALD) as a case study, emphasizing how existing approaches in the literature both build from general approaches used beyond materials science and can be generalized to other materials synthesis techniques. Finally, we provide a practical evaluation framework to evaluate LLMs in the context of materials synthesis
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。