arXiv:2501.10385cs.CYcond-mat.mtrl-sci2025-01被引 8

用大模型驱动的智能助手实现显微镜自动化实验,突破传统流程限制。

Autonomous Microscopy Experiments through Large Language Model Agents

  • 基于大模型代理构建自动化显微实验框架,支持全流程自主操作。
  • 多智能体架构优于单智能体,但主流模型在基础任务上表现不佳。
  • 提示词微调极易影响结果,需谨慎设计以保障实验可靠性。

大语言模型正在重塑材料研究中的自主实验室(SDL),有望加速科学发现。然而,现有实现依赖固定流程,难以模拟专家在动态实验中的应变能力。本文提出人工智能实验助手(AILA),通过大模型驱动代理实现原子力显微镜的自动化操作。同时构建了AFMBench评测基准,全面评估智能体从实验设计到结果分析的完整科学流程。结果显示,当前最优模型在基础任务与协作场景中表现有限;尽管Claude 3.5 Sonnet在材料领域问答任务中表现优异,但在代理任务中却意外表现差,表明领域问答能力无法直接转化为有效代理行为。此外,模型存在偏离指令现象,引发安全对齐担忧。消融实验显示多智能体框架优于单智能体。还发现显著的提示脆弱性:对GPT-4o等强模型,提示结构微调即导致性能剧烈波动。最后,我们评估了AILA在高级实验中的有效性,包括AFM校准、特征检测、力学性质测量、石墨烯层数识别和压头检测。研究强调,在部署AI实验助手前,必须建立严格的基准测试与提示工程策略。

原文摘要 · Abstract (English)

Large language models (LLMs) are revolutionizing self driving laboratories (SDLs) for materials research, promising unprecedented acceleration of scientific discovery. However, current SDL implementations rely on rigid protocols that fail to capture the adaptability and intuition of expert scientists in dynamic experimental settings. We introduce Artificially Intelligent Lab Assistant (AILA), a framework automating atomic force microscopy through LLM driven agents. Further, we develop AFMBench a comprehensive evaluation suite challenging AI agents across the complete scientific workflow from experimental design to results analysis. We find that state of the art models struggle with basic tasks and coordination scenarios. Notably, Claude 3.5 sonnet performs unexpectedly poorly despite excelling in materials domain question answering (QA) benchmarks, revealing that domain specific QA proficiency does not necessarily translate to effective agentic capabilities. Additionally, we observe that LLMs can deviate from instructions, raising safety alignment concerns for SDL applications. Our ablations reveal that multi agent frameworks outperform single-agent architectures. We also observe significant prompt fragility, where slight modifications in prompt structure cause substantial performance variations in capable models like GPT 4o. Finally, we evaluate AILA's effectiveness in increasingly advanced experiments AFM calibration, feature detection, mechanical property measurement, graphene layer counting, and indenter detection. Our findings underscore the necessity for rigorous benchmarking protocols and prompt engineering strategies before deploying AI laboratory assistants in scientific research environments.

自主实验大模型代理原子力显微镜智能科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。