让科研机器人学会评估假设可信度,更聪明地试错。
BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery

- 用贝叶斯方法构建可量化的假设信念状态
- 相同实验次数下,发现效率比传统方法高20%以上
- 适合需要高效探索的自动化科研场景
自主科学发现系统越来越多地使用大语言模型(LLMs)提出新假设,但现有系统主要依赖实验记忆:如高分候选列表或近期试验的启发式总结。本文认为,发现智能体应维护关于假设质量的显式、不确定性感知信念。我们提出 BayesEvolve,一种基于信念的发现框架,将实验证据转化为预测性信念状态,并用此信念指导后续实验。为验证该方法,我们在类BBOB的黑箱优化任务上进行测试,暂未涉及程序与实验室发现领域。结果表明,在固定评估预算下,BayesEvolve 在样本效率上优于基于记忆和档案的 LLM 基线。进一步实验显示,信念状态在保留候选池上具有预测能力;控制性决策规则消融实验表明,带有退火不确定性奖励的信念引导选择更优;且 BayesEvolve 在后期表现出有成效的集中探索,而非无目标的发散尝试。
原文摘要 · Abstract (English)
Autonomous scientific discovery systems increasingly use large language models (LLMs) to propose new hypotheses, but many such systems condition primarily on experimental memory: archives of high-scoring candidates or heuristic summaries of recent trials. We argue that discovery agents should instead maintain explicit, uncertainty-aware beliefs about hypothesis quality. We introduce BayesEvolve, a belief-guided discovery framework that converts experimental evidence into a predictive belief state and uses this belief to guide future experimentation. As a controlled testbed for belief-guided discovery, we evaluate BayesEvolve on shifted BBOB-style black-box optimization tasks, leaving program and laboratory discovery domains to future work. BayesEvolve improves sample efficiency over memory- and archive-guided LLM baselines under a fixed evaluation budget. We further show that the belief state is predictive on held-out candidate pools, that controlled decision-rule ablations favor belief-guided selection with an annealed uncertainty bonus, and that BayesEvolve exhibits productive late-stage concentration rather than unfocused exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。