用大模型+贝叶斯实验设计,高效发现物理机制模型
Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

- 用三层嵌套蒙特卡洛+大模型自动生成新模型假设
- 在3个基准上达到新最优,新基准上仍表现稳健
- 适合需要少样本、抗噪声的科学建模任务
科学的核心目标是从数据中学习机制性或因果性世界模型,这些模型不仅能解释现象,还能回答干预性‘如果……会怎样’的问题。由于被动数据无法识别机制,通常需要实验,但实验成本高,因此需开发数据高效的算法。为此,我们提出模型发现代理(MDA),融合三种技术:一种新型SMC³算法(对模型、参数、潜变量进行三层嵌套序贯蒙特卡洛);一个大语言模型(当当前假设空间不足时,用于提出新模型,实现M-开放贝叶斯推断);以及基于信息价值最大化的实验设计。在三个现有基准—— extit{DPbench}、 extit{CHEMbench} 和 extit{Boxing} 上,MDA 在性能上达到新SOTA。最后,我们引入新的 extit{HHbench} 基准,这是一个更难的随机单神经元电生理数据集,但因具有抗噪声的贝叶斯基础,MDA 依然表现良好。
原文摘要 · Abstract (English)
A primary goal of science is to learn mechanistic or causal world models from data. These models can be used to explain some phenomenon of interest. They also provide the ability to answer interventional ``what if'' questions (i.e., to predict the outcome of an action never taken). Identifying such models usually requires experiments, because passive data leaves the mechanisms unidentified. Since experiments are expensive, we need to develop learning algorithms that are data efficient. We therefore introduce the Model Discovery Agent (MDA), which combines three ingredients: a novel SMC$^3$ algorithm, which uses 3 levels of nested sequential Monte Carlo (over models, parameters, and latents); a large language model (LLM), which is used as a way to propose new models when the current hypothesis space is detected to be insufficient (c.f., M-open Bayesian inference); and an experiment designer based on maximizing the Value of Information. On three existing benchmarks --- \DPbench \citep{wiemann2026discoverphysics}, \CHEMbench \citep{kabra2026autoscilab} and \boxing \citep{gandhi2025boxinggym} --- we show that MDA sets a new SOTA in terms of performance. Finally, we introduce \HHbench, a new stochastic single-neuron electrophysiology benchmark, which is significantly harder than current benchmarks, but on which MDA performs well due to its noise-robust Bayesian foundations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。