用文献支撑推理,生成可验证的科研假说。
HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
- 基于文献构建推理链,小模型学习辨别真假科学论证。
- 假说质量提升,人类评估得分超3.5分,比基线高0.022。
- 适合需要严谨推理的科研人员,尤其关注可验证性。
大型语言模型在跨学科科研创意生成中展现出良好性能。然而,将研究想法转化为可实证验证的具体命题——假设生成,仍较少受到关注。现有方法仅简单引入检索增强,忽视生成过程中的推理逻辑。本文提出HypER(带解释与推理的假说生成),一种训练用于文献引导推理和证据支持型假说生成的小型语言模型。HypER 在多任务设置下训练,能在干扰条件下区分有效与无效的科学推理链条。实验表明,相较于基线模型,HypER 在推理链判别上平均绝对F1提升22%,生成的假说更具证据基础(0.327对比0.305),且经专家评估具备更高可行性与影响力(>3.5分,5点量表)。
原文摘要 · Abstract (English)
Large Language models have demonstrated promising performance in research ideation across scientific domains. Hypothesis development, the process of generating a highly specific declarative statement connecting a research idea with empirical validation, has received relatively less attention. Existing approaches trivially deploy retrieval augmentation and focus only on the quality of the final output ignoring the underlying reasoning process behind ideation. We present $\texttt{HypER}$ ($\textbf{Hyp}$othesis Generation with $\textbf{E}$xplanation and $\textbf{R}$easoning), a small language model (SLM) trained for literature-guided reasoning and evidence-based hypothesis generation. $\texttt{HypER}$ is trained in a multi-task setting to discriminate between valid and invalid scientific reasoning chains in presence of controlled distractions. We find that $\texttt{HypER}$ outperformes the base model, distinguishing valid from invalid reasoning chains (+22\% average absolute F1), generates better evidence-grounded hypotheses (0.327 vs. 0.305 base model) with high feasibility and impact as judged by human experts ($>$3.5 on 5-point Likert scale).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。