arXiv:2507.00310cs.LGcs.AI2025-07NeurIPS被引 30

用贝叶斯惊喜驱动AI自主发现新科学规律。

AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise

  • 以贝叶斯惊喜度为奖励,让AI自主探索假设空间。
  • 在21个真实数据集上,比对手多发现5%-29%的意外结果。
  • 适合追求自动科研发现的学者与研究者使用。

自主科学发现(ASD)的潜力不仅在于回答问题,更在于知道该问什么问题。现有方法多依赖大语言模型(LLM)在目标导向场景中生成假说,但需人类指定研究问题。本文提出AutoDiscovery——一种开放式的科学发现方法,通过贝叶斯惊喜度驱动探索。该方法量化了LLM从先验信念到实验后后验信念的认知转变。为高效搜索嵌套假设空间,采用蒙特卡洛树搜索(MCTS)结合渐进式扩宽,以惊喜度为奖励函数。我们在21个真实世界数据集(涵盖生物、经济、金融、行为科学等)上评估该方法。结果表明,在固定预算下,AutoDiscovery比竞争对手多产生5%-29%被LLM判定为意外的发现。人工评估显示,系统发现中有三分之二对领域专家而言也具有意外性,标志着向构建开放式自主科学发现系统迈出重要一步。

原文摘要 · Abstract (English)

The promise of autonomous scientific discovery (ASD) hinges not only on answering questions, but also on knowing which questions to ask. Most recent works in ASD explore the use of large language models (LLMs) in goal-driven settings, relying on human-specified research questions to guide hypothesis generation. However, scientific discovery may be accelerated further by allowing the AI system to drive exploration by its own criteria. The few existing approaches in open-ended ASD select hypotheses based on diversity heuristics or subjective proxies for human interestingness, but the former struggles to meaningfully navigate the typically vast hypothesis space, and the latter suffers from imprecise definitions. This paper presents AutoDiscovery -- a method for open-ended ASD that instead drives scientific exploration using Bayesian surprise. Here, we quantify the epistemic shift from the LLM's prior beliefs about a hypothesis to its posterior beliefs after gathering experimental results. To efficiently explore the space of nested hypotheses, our method employs a Monte Carlo tree search (MCTS) strategy with progressive widening using surprisal as the reward function. We evaluate AutoDiscovery in the setting of data-driven discovery across 21 real-world datasets spanning domains such as biology, economics, finance, and behavioral science. Our results demonstrate that under a fixed budget, AutoDiscovery substantially outperforms competitors by producing 5-29% more discoveries deemed surprising by the LLM. Our human evaluation further reveals that two-thirds of discoveries made by our system are surprising to domain experts as well, suggesting this is an important step towards building open-ended ASD systems.

科学发现贝叶斯大模型自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。