用多个AI代理自动优化扫描透射电镜成像参数,提升成功率和效率。
PEAR: A Robust and Flexible Automation Framework for Ptychography Enabled by Multiple Large Language Model Agents
- 设计多智能体系统,分工完成知识检索、代码生成与参数推荐。
- 在小模型(LLaMA 3.1 8B)下仍实现高成功率,实验成功率达90%以上。
- 支持本地知识库定制,适合科研与工业场景灵活部署。
扫描透射电镜(Ptychography)是一种先进的计算成像技术,广泛应用于物理、化学、生物及材料科学等领域,也用于半导体表征等工业场景。高质量成像需同时优化大量实验与算法参数,传统方法依赖试错,效率低且易受人为偏见影响。本文提出‘扫描透射电镜实验与分析机器人’(PEAR),利用大语言模型(LLMs)自动化数据处理流程。为确保鲁棒性与准确性,PEAR采用多智能体架构,分别执行知识检索、代码生成、参数建议与图像推理任务。实验表明,即使使用较小的开源模型(如LLaMA 3.1 8B),PEAR的多智能体设计也能显著提升工作流成功率(>90%)。该框架支持不同自动化层级,并可集成自定义本地知识库,具备跨研究环境的灵活性与适应性。
原文摘要 · Abstract (English)
Ptychography is an advanced computational imaging technique in X-ray and electron microscopy. It has been widely adopted across scientific research fields, including physics, chemistry, biology, and materials science, as well as in industrial applications such as semiconductor characterization. In practice, obtaining high-quality ptychographic images requires simultaneous optimization of numerous experimental and algorithmic parameters. Traditionally, parameter selection often relies on trial and error, leading to low-throughput workflows and potential human bias. In this work, we develop the "Ptychographic Experiment and Analysis Robot" (PEAR), a framework that leverages large language models (LLMs) to automate data analysis in ptychography. To ensure high robustness and accuracy, PEAR employs multiple LLM agents for tasks including knowledge retrieval, code generation, parameter recommendation, and image reasoning. Our study demonstrates that PEAR's multi-agent design significantly improves the workflow success rate, even with smaller open-weight models such as LLaMA 3.1 8B. PEAR also supports various automation levels and is designed to work with customized local knowledge bases, ensuring flexibility and adaptability across different research environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。