arXiv:2603.26718cs.CYcs.AI2026-03被引 1

构建科学多智能体系统评估框架,解决推理与检索混淆、数据污染等难题。

Toward Evaluation Frameworks for Multi-Agent Scientific AI Systems

  • 提出抗污染任务构建策略,避免模型训练数据泄露。
  • 设计可扩展的任务家族,支持多轮交互测试系统能力。
  • 结合科研人员访谈,指导评估方法贴近真实科研流程。

我们分析了评测科学(多智能体)系统面临的挑战,包括难以区分推理与检索、数据/模型污染风险、新颖研究问题缺乏可靠真值、工具使用带来的复杂性,以及因知识库持续更新导致的复现困难。文章讨论了构建抗污染问题、生成可扩展任务族的策略,并强调需通过多轮交互评估来更真实反映科研实践。作为初步可行性验证,我们构建了一个包含新研究构想的数据集,用于测试系统的泛化性能。此外,还基于对多位量子科学领域研究人员和工程师的访谈,探讨科学家对AI系统的期望,这些期望应指导评估方法的设计。

原文摘要 · Abstract (English)

We analyze the challenges of benchmarking scientific (multi)-agentic systems, including the difficulty of distinguishing reasoning from retrieval, the risks of data/model contamination, the lack of reliable ground truth for novel research problems, the complications introduced by tool use, and the replication challenges due to the continuously changing/updating knowledge base. We discuss strategies for constructing contamination-resistant problems, generating scalable families of tasks, and the need for evaluating systems through multi-turn interactions that better reflect real scientific practice. As an early feasibility test, we demonstrate how to construct a dataset of novel research ideas to test the out-of-sample performance of our system. We also discuss the results of interviews with several researchers and engineers working in quantum science. Through those interviews, we examine how scientists expect to interact with AI systems and how these expectations should shape evaluation methods.

多智能体科学计算评估框架量子科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。