arXiv:2607.27556cs.AIcs.MA2026-07

提出FEV框架,用可追溯流程评估生物信息学智能体的科学可信度。

Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

论文配图:Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
图 1 · 摘自论文原文
  • 以可检查的工作流轨迹为核心,分离功能、证据与验证三要素。
  • 分析109个系统和28个评测资源,发现流程可复现性滞后于执行能力。
  • 适合关注科研可信度与工作流透明性的生物信息学研究者使用。

大型语言模型代理在生物分析中日益承担规划、执行与解释任务,但流畅回应、成功工具调用及基准表现不足以证明其科学可信性。现有综述多按应用、架构或智能体能力分类,却未协同量化代理工作流的责任性。本文提出功能-证据-验证(FEV)框架,将可检查的工作流轨迹作为核心分析单元,分离已实现的操作、可追溯的支持依据以及特定场景的验证。基于该框架,我们梳理了109个智能体或类智能体系统及28个评测资源,涵盖基因组学、单细胞与空间组学、蛋白质科学、药物发现、计算病理学和通用生物信息自动化等领域的128篇论文。结果显示,规划与工具执行进展迅速,但可复现性、溯源性、稳健科学评估、外部验证和前瞻性实证测试仍显著滞后。因此主张以工作流正确性而非最终答案正确性来评价智能体生物信息学系统。FEV为系统比较和设计透明、可审计、科学可问责的工作流提供了实践基础。

原文摘要 · Abstract (English)

Large language model agents increasingly plan, execute, and interpret biological analyses, yet fluent responses, successful tool calls, and benchmark performance alone do not establish scientific credibility. Existing reviews primarily organize biological agents by application, architecture, and agentic capability, but do not jointly operationalize the accountability of agent-generated workflows. We address this gap by treating the inspectable workflow trajectory, rather than architecture or final output alone, as the primary unit of analysis. We introduce the Function--Evidence--Validation (FEV) framework, which separates demonstrated workflow operations, traceable support for actions and claims, and use-case-specific validation. Using FEV, we map 109 agentic or agent-adjacent systems and 28 benchmark or evaluation resources, representing 128 unique publications across genomics, single-cell and spatial omics, protein science, drug discovery, computational pathology, and general bioinformatics automation. Across domains, planning and tool-mediated execution have advanced more rapidly than replayability, provenance, robust scientific assessment, external validation, and prospective empirical testing. We therefore argue that agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone. FEV provides a practical basis for comparing systems and designing transparent, auditable, and scientifically accountable bioinformatics workflows.

智能体评估生物信息学可解释性工作流验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。