arXiv:2505.14599cs.CLcs.AI2025-05IJCAI被引 13

评测大模型生成科学假说的真假能力,发现其易幻觉,提出新检测工具。

Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

  • 构建基准TruthHypo与知识检测器KnowHD评估假说可靠性。
  • 大模型生成假说多不真实,且存在严重幻觉现象。
  • KnowHD能有效筛选可信假说,助科研人员加速发现。

大语言模型(LLMs)在生物医学等科学领域展现出巨大潜力,尤其在假设生成方面:可分析海量文献、识别模式并提出研究方向。然而,验证生成假说的真实性常需大量时间和资源。此外,模型幻觉问题导致生成看似合理实则错误的假说,影响其可靠性。为此,我们提出TruthHypo基准,用于系统评估LLMs生成科学假说的能力,并开发基于知识的幻觉检测器KnowHD,评估假说是否扎根于现有知识。实验表明,LLMs难以生成真实假说;通过分析推理过程中的幻觉,我们证明KnowHD提供的置信度分数可有效过滤出可信假说。人工评估进一步验证了KnowHD在识别真实假说和加速科学发现方面的价值。数据与代码已开源:https://github.com/Teddy-XiongGZ/TruthHypo。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown significant potential in scientific disciplines such as biomedicine, particularly in hypothesis generation, where they can analyze vast literature, identify patterns, and suggest research directions. However, a key challenge lies in evaluating the truthfulness of generated hypotheses, as verifying their accuracy often requires substantial time and resources. Additionally, the hallucination problem in LLMs can lead to the generation of hypotheses that appear plausible but are ultimately incorrect, undermining their reliability. To facilitate the systematic study of these challenges, we introduce TruthHypo, a benchmark for assessing the capabilities of LLMs in generating truthful scientific hypotheses, and KnowHD, a knowledge-based hallucination detector to evaluate how well hypotheses are grounded in existing knowledge. Our results show that LLMs struggle to generate truthful hypotheses. By analyzing hallucinations in reasoning steps, we demonstrate that the groundedness scores provided by KnowHD serve as an effective metric for filtering truthful hypotheses from the diverse outputs of LLMs. Human evaluations further validate the utility of KnowHD in identifying truthful hypotheses and accelerating scientific discovery. Our data and source code are available at https://github.com/Teddy-XiongGZ/TruthHypo.

科学假说幻觉检测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。