用视觉语言模型自动完成显微镜实验,降低操作门槛。
EAA: Automating materials characterization with vision language model agents
- 基于多模态推理与工具增强的智能代理系统
- 实现全自动对焦、自然语言搜索特征等流程
- 适合科研人员快速上手复杂仪器操作
我们提出实验自动化代理(EAA),一个由视觉语言模型驱动的智能体系统,用于自动化复杂的显微镜实验流程。EAA结合多模态推理、工具增强动作及可选长期记忆,支持完全自主或用户交互式测量。基于灵活的任务管理架构,系统可实现从全代理驱动到嵌入局部LLM查询的逻辑化流程。EAA还提供现代化工具生态,兼容Model Context Protocol(MCP),实现仪器控制工具在应用间的双向调用。我们在先进光子源成像束线验证了EAA,包括自动聚焦透镜、自然语言描述的特征搜索和交互式数据采集。结果表明,具备视觉能力的智能体能提升束线效率,减轻操作负担,并降低用户使用门槛。
原文摘要 · Abstract (English)
We present Experiment Automation Agents (EAA), a vision-language-model-driven agentic system designed to automate complex experimental microscopy workflows. EAA integrates multimodal reasoning, tool-augmented action, and optional long-term memory to support both autonomous procedures and interactive user-guided measurements. Built on a flexible task-manager architecture, the system enables workflows ranging from fully agent-driven automation to logic-defined routines that embed localized LLM queries. EAA further provides a modern tool ecosystem with two-way compatibility for Model Context Protocol (MCP), allowing instrument-control tools to be consumed or served across applications. We demonstrate EAA at an imaging beamline at the Advanced Photon Source, including automated zone plate focusing, natural language-described feature search, and interactive data acquisition. These results illustrate how vision-capable agents can enhance beamline efficiency, reduce operational burden, and lower the expertise barrier for users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。