arXiv:2512.04854cs.AI2025-12

评估AI在生物医学研究中作科研伙伴的能力,而非仅执行任务。

From Task Executors to Research Partners: Evaluating AI Co-Pilots Through Workflow Integration in Biomedical Research

  • 提出以工作流整合为核心的评估框架
  • 发现现有基准只测单一能力,不反映真实协作
  • 适合关注AI科研助手实用性的研究者

人工智能系统在生物医学研究中日益普及,但现有评估框架可能无法有效衡量其作为研究合作者的实际效能。本快速综述分析了2018年1月1日至2025年10月31日期间三大数据库及两个预印本平台,共识别出14项评估AI在文献理解、实验设计和假说生成方面能力的基准。结果显示,所有现有基准均仅评估孤立组件能力,如数据分析质量、假说有效性与实验方案设计。然而,真实的科研协作需要跨多轮会话的集成工作流,具备上下文记忆、自适应对话与约束传播能力。这一差距表明,即使在组件基准上表现优异的系统,在实际科研协作中也可能失效。为此,本文提出一种过程导向的评估框架,涵盖当前基准缺失的四个关键维度:对话质量、工作流编排、会话连续性与研究人员体验。这些维度对评估AI作为研究协作者至关重要,而非仅作为任务执行器。

原文摘要 · Abstract (English)

Artificial intelligence systems are increasingly deployed in biomedical research. However, current evaluation frameworks may inadequately assess their effectiveness as research collaborators. This rapid review examines benchmarking practices for AI systems in preclinical biomedical research. Three major databases and two preprint servers were searched from January 1, 2018 to October 31, 2025, identifying 14 benchmarks that assess AI capabilities in literature understanding, experimental design, and hypothesis generation. The results revealed that all current benchmarks assess isolated component capabilities, including data analysis quality, hypothesis validity, and experimental protocol design. However, authentic research collaboration requires integrated workflows spanning multiple sessions, with contextual memory, adaptive dialogue, and constraint propagation. This gap implies that systems excelling on component benchmarks may fail as practical research co-pilots. A process-oriented evaluation framework is proposed that addresses four critical dimensions absent from current benchmarks: dialogue quality, workflow orchestration, session continuity, and researcher experience. These dimensions are essential for evaluating AI systems as research co-pilots rather than as isolated task executors.

AI科研评估框架协同工作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。