arXiv:2505.16100cs.AIcs.CL2025-05被引 8

构建1029个生物医学假说验证任务,评测AI在真实科研中的推理能力。

BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

  • 从300+论文中提取真实研究流程,构建带证据的假说任务
  • 包含可验证与不可验证假说,覆盖真实科研复杂性
  • 支持多维度评估:判断准确率、推理逻辑、代码可执行性

生物医学研究中验证科学假说是核心挑战,而现有AI代理因真实数据分析与证据解读的复杂性仍难胜任。本文提出BioDSA-1K,一个包含1,029个以假说为中心的任务和1,177个分析计划的基准,数据源自300余篇已发表研究,反映真实科研工作流结构与推理模式。每个任务包含从原研究结论提炼的肯定式假说及基于实证数据表的支持证据,虽与已发表结论一致,但可通过标准统计或机器学习方法重新检验。基准支持四维评估:(1)假说判断准确率,(2)证据与结论的一致性,(3)推理过程正确性,(4)AI生成代码的可执行性。特别地,包含非可验证假说——即数据不足以支持或反驳的案例,体现现实科研中常见却未被充分研究的情形。本基准旨在推动通用、可信生物医学发现AI代理的发展。

原文摘要 · Abstract (English)

Validating scientific hypotheses is a central challenge in biomedical research, and remains difficult for artificial intelligence (AI) agents due to the complexity of real-world data analysis and evidence interpretation. In this work, we present BioDSA-1K, a benchmark designed to evaluate AI agents on realistic, data-driven biomedical hypothesis validation tasks. BioDSA-1K consists of 1,029 hypothesis-centric tasks paired with 1,177 analysis plans, curated from over 300 published biomedical studies to reflect the structure and reasoning found in authentic research workflows. Each task includes a structured hypothesis derived from the original study's conclusions, expressed in the affirmative to reflect the language of scientific reporting, and one or more pieces of supporting evidence grounded in empirical data tables. While these hypotheses mirror published claims, they remain testable using standard statistical or machine learning methods. The benchmark enables evaluation along four axes: (1) hypothesis decision accuracy, (2) alignment between evidence and conclusion, (3) correctness of the reasoning process, and (4) executability of the AI-generated analysis code. Importantly, BioDSA-1K includes non-verifiable hypotheses: cases where the available data are insufficient to support or refute a claim, reflecting a common yet underexplored scenario in real-world science. We propose BioDSA-1K as a foundation for building and evaluating generalizable, trustworthy AI agents for biomedical discovery.

AI科研假说验证生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。