用模型自身置信度评分,比让模型比较优劣更有效评估科学假说。
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

- 基于语言模型输出的对数能量值,衡量假说内在可信度。
- 在12个学科1323篇论文上,最高准确率达53.1%。
- 适合需要客观筛选新假设的研究者使用。
大型语言模型(LLMs)在科学假说生成中应用日益广泛,但生成假说的评估仍面临挑战。现有方法常依赖模型作为评判者或语义相似性,易偏好熟悉观点而非新颖假说。本文提出一种基于对数能量的内在评分方法,利用语言模型自身的置信度进行评价,而非依赖对比判断。我们在12个学科的1,323篇论文上进行了基准测试,每篇论文配有一个正确假说和十五个错误替代项。结果表明,内在评分在所有评分器上平均达到33.0% Hit@1,远超提示式列表排名的16.6%。最强配置——一个10亿参数模型结合对数能量评分,达到53.1%的最高准确率(在14种模型-评分组合中后验选出)。研究显示,基于置信度的方法在科学假说评估中具有潜力,并为可信人工智能驱动的科学发现提供了新方向。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similarity, which can favor familiar ideas over novel ones. We propose a logit-based energy scoring method that evaluates hypotheses using a language model's intrinsic confidence rather than comparative judgment. We benchmarked seven language models on 1,323 papers across 12 disciplines. Each paper was paired with its hypothesis and fifteen incorrect alternatives. Intrinsic scoring reached 33.0% Hit@1 pooled across both scorers, compared with 16.6% for prompted listwise ranking. The strongest configuration, a 1-billion-parameter model using logit-based energy scoring, reached 53.1%, though this was the maximum across 14 model-by-scorer combinations selected post hoc. Overall, intrinsic model confidence shows potential for scientific hypothesis evaluation. This study also motivates future research on confidence-based methods for trustworthy AI-enabled scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。