让AI生成可验证且有文献依据的科研假设,还能根据实验结果自动优化。
HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation
- 模仿人类研究者流程,从文献中发现趋势并设计可验证假说。
- 生成假说在可行性与文献依据上显著提升,实测成功率高出28%。
- 能根据实验反馈动态调整,适合需要持续迭代的科研项目。
尽管自动化科学发现(ASD)因大模型兴起而备受关注,但现有工具仍难以生成既可验证又扎根于科学文献的假说,且缺乏对先前实验结果的适应能力。我们提出HARPA框架,借鉴人类研究人员的思维流程:首先通过文献挖掘识别新兴研究趋势,再探索假说设计空间,最后通过定位研究空白和论证设计选择,收敛到精确、可验证的假说。评估显示,HARPA生成的研究提案在多数定性维度(如具体性、新颖性、整体质量)上媲美强基线AI研究员,但在可行性(+0.78, p<0.05)与文献根基性(+0.85, p<0.01)上显著提升(10分量表)。在与ASD代理CodeScientist测试中,HARPA实现20次成功执行(40次中),失败仅16次,远优于对照组(11次成功,21次失败),表明专家可行性判断与实际执行成功率高度一致。此外,通过学习基于过往实验结果的奖励模型,HARPA的假说评分相较未训练基线提升约28%绝对值。这些方法推动了人工智能驱动科学发现的发展。
原文摘要 · Abstract (English)
While there has been a surge of interest in automated scientific discovery (ASD), especially with the emergence of LLMs, it remains challenging for tools to generate hypotheses that are both testable and grounded in the scientific literature. Additionally, existing ideation tools are not adaptive to prior experimental outcomes. We developed HARPA to address these challenges by incorporating the ideation workflow inspired by human researchers. HARPA first identifies emerging research trends through literature mining, then explores hypothesis design spaces, and finally converges on precise, testable hypotheses by pinpointing research gaps and justifying design choices. Our evaluations show that HARPA-generated hypothesis-driven research proposals perform comparably to a strong baseline AI-researcher across most qualitative dimensions (e.g., specificity, novelty, overall quality), but achieve significant gains in feasibility(+0.78, p$<0.05$, bootstrap) and groundedness (+0.85, p$<0.01$, bootstrap) on a 10-point Likert scale. When tested with the ASD agent (CodeScientist), HARPA produced more successful executions (20 vs. 11 out of 40) and fewer failures (16 vs. 21 out of 40), showing that expert feasibility judgments track with actual execution success. Furthermore, to simulate how researchers continuously refine their understanding of what hypotheses are both testable and potentially interesting from experience, HARPA learns a reward model that scores new hypotheses based on prior experimental outcomes, achieving approx. a 28\% absolute gain over HARPA's untrained baseline scorer. Together, these methods represent a step forward in the field of AI-driven scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。