自动从文献构建贝叶斯校准的先验分布,提升科学建模可解释性。
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

- 用多智能体系统搜索文献,按领域相关性加权并拟合概率分布
- 在24个参数上生成可溯源先验,比单提示基线更少错误输出
- 本地运行保障数据隐私,适合有领域知识但缺统计经验的研究者
过程模型的贝叶斯校准需要为每个参数设定先验分布。尽管已有多年方法研究,研究人员仍普遍使用均匀先验。主要原因是基于文献构建信息性先验耗时且需跨领域与统计专长。我们提出Distribird,一个自动化工具,输入参数名、物理描述和领域背景后,通过多智能体流程检索文献,按领域相关性提取并加权报告值,再以AIC准则选择概率分布拟合。若无文献支持,则采用合理的非信息性替代方案,并明确标注依据与置信度。该工具适用于具有物理可解释参数且领域知识存在于已发表文献中的场景。我们在10个科学领域共24个参数上评估,对比三个开源大模型(Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B)与单提示基线。结果表明,完整流程在先验质量上匹配基线;所有先验均可追溯至具体文献与数值;内置有效性层拒绝超出范围请求,而单提示基线在30个模型-参数对中11次返回自信但无根据的先验;所有语言模型调用均本地执行,仅生成的搜索词传至公开文献数据库,不泄露参数描述或未发表建模细节。我们认为这些特性对科学应用比微小点估计精度提升更重要。
原文摘要 · Abstract (English)
Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present Distribird, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system falls back to sensible uninformative alternatives, and clearly reports both the evidence behind and the confidence level of every prior it produces. It is designed for the problems where the models have physically interpretable parameters, where domain knowledge exists in the published literature. We evaluate the tool on 24~parameters across 10 scientific domains comparing three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) with a single-prompt LLM baseline. On prior quality the full pipeline matches this baseline. Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30~model-parameter cases; and every language-model call runs locally, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). For scientific use, we argue these properties matter more than a marginal improvement in point-estimate accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。