评测AI代理从文献中自动整理果蝇知识的全流程能力
FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases
- 设计端到端任务框架,让AI代理从论文中提取基因功能与命名历史
- 多智能体架构表现最优,但大模型扩展效果递减,仍有巨大提升空间
- 适合研究科学推理、知识库构建与AI辅助科研的学者参考
科学知识库通过将原始文献中的发现结构化为可查询格式,加速科学发现。维护这些资源需专家搜索论文、整合证据并生成基于本体的注释,而现有基准仅关注命名实体识别或关系抽取等单一任务,无法覆盖完整流程。本文提出FlyBench,用于评估AI代理在果蝇科学文献中进行端到端本体注释的能力。给定一个基因符号,代理需从16,898篇全文论文中检索并阅读,生成包含功能、表达模式及历史同义词的结构化注释。该基准涵盖来自FlyBase的100个基因、7,397条专家标注数据。我们评估了四种基线代理架构:记忆型、固定流水线、单智能体和多智能体。结果表明,架构选择显著影响性能,多智能体优于简单方案,但扩大模型规模带来的收益递减。分析显示,代理主要依赖检索验证已有知识,而非发现新信息。我们希望FlyBench能推动检索增强型科学推理的发展,该能力在多个科学领域均有广泛应用前景。
原文摘要 · Abstract (English)
Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert curators to search relevant papers, reconcile evidence across documents, and produce ontology-grounded annotations - a workflow that existing benchmarks, focused on isolated subtasks like named entity recognition or relation extraction, do not capture. We present FlyBench to evaluate AI agents on end-to-end agentic ontology curation from scientific literature. Given only a gene symbol, agents must search and read from a corpus of 16,898 full-text papers to produce structured annotations: Gene Ontology terms describing function, expression patterns, and historical synonyms linking decades of nomenclature. The benchmark includes 7,397 expert-curated annotations across 100 genes drawn from FlyBase, the Drosophila (fruit fly) knowledge base. We evaluate four baseline agent architectures: memorization, fixed pipeline, single-agent, and multi-agent. We find that architectural choices significantly impact performance, with multi-agent designs outperforming simpler alternatives, yet scaling backbone models yields diminishing returns. All baselines leave substantial room for improvement. Our analysis surfaces several findings to guide future development; for example, agents primarily use retrieval to confirm parametric knowledge rather than discover new information. We hope FlyBench will drive progress on retrieval-augmented scientific reasoning, a capability with broad applications across scientific domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。