arXiv:2606.07718cs.AIcs.CV2026-06

评估通用编程智能体在果蝇光遗传学数据发现流程中的表现

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

论文配图:A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline
图 1 · 摘自论文原文
  • 用真实科研流程测试智能体,任务规模远超现有基准
  • 智能体可完成部分流程阶段,但整体端到端执行仍无法实现
  • 缺乏预设标准时,智能体难以自主判断结果正确性

代理型AI为自动化科学研发流程中的软件开发瓶颈提供了前景,尤其适用于领域专家需数天至数月构建、且正确性和鲁棒性远高于实现细节的环节。我们对通用编程智能体在果蝇光遗传学数据到发现的全流程中进行了实证研究。评估任务和数据集规模显著大于现有基准,评价标准基于领域专家标准。结果表明,智能体可解决多个流程阶段,说明阶段级自动化是可行的。通过分析代码迭代过程,发现当缺乏预定义标准时,智能体依赖科学判断评估当前方案,表现不佳;虽会尝试查看中间输出进行自评,但大多无法正确解读或采取有效行动。端到端流程需串联各阶段成功,超出当前智能体能力。我们识别出现有基准中缺失的关键挑战,包括计算资源管理及对未见数据的泛化能力。最后,提出构建科学任务与开放问题严谨评估标准的原则。

原文摘要 · Abstract (English)

Agentic AI offers a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build and where correctness and robustness matter more than implementation details. We present an empirical study of general-purpose coding agents on a fly optogenetics data-to-discovery pipeline. We assess agents on tasks and datasets substantially larger than existing benchmarks and evaluation criteria grounded in domain expert standards. We show that agents can solve several pipeline stages, suggesting stage-level automation is tractable. By analyzing agents' code iterations, we show they struggle most without a pre-defined criterion, when they must instead use their scientific judgment to assess their current solution. Mirroring scientists, they sometimes attempt visual inspection of intermediate outputs for self-evaluation, but largely fail to interpret what they see or act on it appropriately. Solving the end-to-end pipeline requires stringing together successes across all stages, which is beyond agents' current abilities. We identify challenges largely absent from existing benchmarks, including computational resource management and generalization to held-out data. Finally, we distill principles for constructing scientific tasks and rigorous evaluation criteria for open-ended problems.

AI代理科学发现自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。