arXiv:2608.05266cs.AIcond-mat.mtrl-sci2026-08

测试大模型显微镜代理性能,发现基准测试能选优但难预测新任务表现。

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

论文配图:Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks
图 1 · 摘自论文原文
  • 构建多代理架构与日志框架,评估53项测试中105种配置
  • 不同配置在延迟、耗时、成本上差异显著,但无法可靠预测新任务表现
  • 适合做系统优化和故障诊断,不适用于通用配置推荐

大语言模型代理正被用于控制显微镜等科学仪器,但其系统设计尚无统一范式。本文构建了基准测试与追踪日志框架,评估了一至三代理图拓扑、五种LLM、RAG参数及运行约束,在53项显微镜任务中测试了105种代理配置,共完成1,949次测试运行与49,109次RAG检索。结果表明,不同配置在延迟、令牌消耗、成本和失败模式上存在明显差异。然而,基于架构与测试结果训练的替代模型无法可靠预测代理在未见任务上的表现。这说明当前基准仅适用于资格认证、回归测试、故障诊断与直接对比,难以支撑任务无关的全局配置模型。

原文摘要 · Abstract (English)

Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent and there are few well-established paradigms for how to engineer an agentic system. There are many choices to make when designing a microscopy agent, including the choice of LLM, the number of agents to use, agent responsibilities and delegation rules, retrieval-augmented generation parameters, and more. When designing and optimizing an agentic microscope controller, researchers not only want to ensure that the agent can correctly perform known tasks but also that the agent can generalize to new tasks that it has not encountered before. In this study, we develop a benchmark and trace-logging framework that reveals a) how different choices of agent architecture impact performance at microscopy tasks and b) the limitations of benchmarks for predicting if a particular agent will perform well on unseen microscopy tasks. The framework was used to evaluate one-, two-, and three-agent graph topologies, five LLMs, RAG and context parameters, and operational constraints across 53 microscopy benchmark tests. In total, 105 agent configurations, 1,949 individual test runs, and 49,109 RAG retrievals were recorded. Direct comparisons showed clear differences in latency, token use, cost, and failure mode between configurations. However, surrogate models trained on agent architecture and test results did not reliably predict an agent's performance on new, unseen tasks. These results show that these benchmarks are useful for qualification, regression testing, diagnosis, and direct comparison, but the current heterogeneous test suite does not support a task-independent global configuration model.

智能显微镜LLM代理基准测试泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。