arXiv:2503.21248cs.CLcs.AI2025-03ACL被引 57

首个科学发现基准,评估大模型生成高质量研究假设能力

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

  • 基于灵感分解法构建科学发现任务框架
  • 大模型在跨领域灵感检索上表现优异,能发现新颖知识关联
  • 专为2024年后论文设计,避免数据污染,支持持续更新

大型语言模型(LLMs)在辅助科研方面展现出潜力,但其生成高质量研究假说的能力尚未得到充分评估,原因在于缺乏专用基准。为此,我们提出了首个大规模基准,用于评估LLMs在科学发现的三个关键子任务上的表现:灵感检索、假说构建和假说排序。其中,这三个子任务的完备解决可完全实现整体发现目标。我们开发了一个基于LLM的自动化框架,从12个学科领域的论文中提取关键要素——研究问题、背景综述、灵感来源和假说,并通过专家验证确认其准确性。为防止数据污染,我们仅使用2024年及之后发表的文献,确保与模型预训练数据重叠极小;该自动化框架还能随预训练截止日期更新而自动提取更近期论文,支持基准的可扩展、无污染持续更新。评估结果显示,在多个学科中,LLMs在灵感检索这一分布外任务上表现突出,表明其具备发现新颖知识关联的能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark. To address this gap, we introduce the first large-scale benchmark for evaluating LLMs on a sufficient set of scientific discovery sub-tasks-inspiration retrieval, hypothesis composition, and hypothesis ranking-where sufficient means that perfectly solving these sub-tasks perfectly solves the overall discovery task. We develop an automated LLM-based framework that extracts critical components-research questions, background surveys, inspirations, and hypotheses-from papers across 12 disciplines, with expert validation confirming its accuracy. To prevent data contamination, we focus exclusively on publications from 2024 onward, ensuring minimal overlap with LLM pretraining data; our automated framework further enables automatic extraction of even more recent papers as LLM pretraining cutoffs advance, supporting scalable and contamination-free automatic renewal of this discovery benchmark. Our evaluation shows that, across disciplines, LLMs excel at inspiration retrieval-an out-of-distribution task-suggesting their ability to surface novel knowledge associations.

科学发现大模型评估灵感检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。