arXiv:2506.21964stat.MEcs.AI2025-06

用大模型生成贝叶斯分析的先验分布,提升效率与客观性。

Using Large Language Models to Suggest Informative Prior Distributions in Bayesian Statistics

  • 通过精心设计提示词,让大模型推荐并自检先验分布。
  • 大模型准确识别变量关联方向,但中等先验常过自信。
  • Claude在弱先验上表现更优,避免无意义的零均值假设。

贝叶斯统计中先验分布的选择既困难又主观。本文探讨使用大语言模型(LLMs)生成基于知识的有信息先验。我们设计了包含建议、验证与反思的复杂提示,评估了Claude Opus、Gemini 2.5 Pro和ChatGPT-4o-mini在两个真实数据集(心脏病风险与混凝土强度)上的表现。所有模型均正确识别出所有关联方向(如男性心脏病风险更高)。先验质量通过其与最大似然估计分布的KL散度衡量。模型提出中等和弱信息先验。中等先验常过于自信,与数据不匹配;而弱先验中,ChatGPT与Gemini倾向默认零均值,导致过度模糊,相比之下Claude未出现此问题,表现更优。大模型在识别正确关系方面展现出巨大潜力,但核心挑战仍在于调节先验宽度,避免过自信或低估。

原文摘要 · Abstract (English)

Selecting prior distributions in Bayesian statistics is challenging, resource-intensive, and subjective. We analyze using large-language models (LLMs) to suggest suitable, knowledge-based informative priors. We developed an extensive prompt asking LLMs not only to suggest priors but also to verify and reflect on their choices. We evaluated Claude Opus, Gemini 2.5 Pro, and ChatGPT-4o-mini on two real datasets: heart disease risk and concrete strength. All LLMs correctly identified the direction for all associations (e.g., that heart disease risk is higher for males). The quality of suggested priors was measured by their Kullback-Leibler divergence from the maximum likelihood estimator's distribution. The LLMs suggested both moderately and weakly informative priors. The moderate priors were often overconfident, resulting in distributions misaligned with the data. In our experiments, Claude and Gemini provided better priors than ChatGPT. For weakly informative priors, a key performance difference emerged: ChatGPT and Gemini defaulted to an "unnecessarily vague" mean of 0, while Claude did not, demonstrating a significant advantage. The ability of LLMs to identify correct associations shows their great potential as an efficient, objective method for developing informative priors. However, the primary challenge remains in calibrating the width of these priors to avoid over- and under-confidence.

贝叶斯统计大模型应用先验分布AI辅助建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。