arXiv:2606.29182cs.AIcs.CL2026-06

让大模型动态更新科学信念,提升持续发现能力

Evidence-Informed LLM Beliefs for Continual Scientific Discovery

  • 用过往假设证据动态更新模型先验,实现非静态惊奇度计算
  • 新方法使累积惊奇度平均提升30.62%,识别37.5%的虚假高分假设
  • 适合做持续性科学探索、需避免重复与鼓励多样性的研究者

以大语言模型(LLMs)进行开放式科学发现,正形成一个长期循环:提出假设并验证,由奖励信号引导下一步探索。近期代表工作AutoDiscovery采用“贝叶斯惊奇度”作为发现指标和搜索奖励,即模型在观察假设证据后信念的变化。我们发现,AutoDiscovery将惊奇度视为静态量,但人类推理中的惊奇度是随经验演变的动态过程,而这是持续科学发现的前提。为此,我们提出证据驱动的LLM信念机制:利用先前假设的证据更新先验,从而计算新假设的非静态惊奇度。比较多种上下文信念更新方式后发现,基于嵌入的检索增强生成在预测最终后验方面表现最优,能识别出37.5%的静态惊奇度为虚假。进而,我们改进搜索策略,通过信念更新过滤和多样性最大化,避免虚假奖励与冗余探索。在五个发现领域中,新方法使累积非静态惊奇度平均提升30.62%,表明持续科学发现不仅需要更精准的信念测量,还需能规避重复、促进多样性的搜索机制。

原文摘要 · Abstract (English)

Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides which hypotheses to test next. A notable recent example is AutoDiscovery, which uses "Bayesian surprise" - the belief shift an LLM undergoes after observing evidence for a hypothesis - as both a discovery metric and a reward for search. We first observe that AutoDiscovery treats surprisal as a static quantity, while surprisal in human reasoning is non-stationary - it is defined relative to beliefs that evolve with experience, a prerequisite for continual scientific discovery. We address this mismatch with evidence-informed LLM beliefs: priors updated with evidence from previous hypotheses to compute non-stationary surprisal for new hypotheses. We compare in-context belief-updating mechanisms and find that embedding-based retrieval-augmented generation over prior discoveries best anticipates eventual posteriors, identifying 37.5% of static surprisals as spurious. We then modify search to avoid these spurious rewards and prioritize hypotheses that remain surprising under non-stationary beliefs. Concretely, we introduce two complementary changes to the original search procedure: belief-update filtering and diversity maximization. Across five discovery domains, our method increases accumulated non-stationary surprisal by 30.62% on average compared to the original search procedure, demonstrating that continual scientific discovery with LLMs requires not only better belief measurement but also search procedures that avoid redundancy and encourage diversity.

科学发现大模型动态信念搜索优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。