arXiv:2604.01896cs.AI2026-04

大模型估不准概率,越大越准,但多想没用,还得调校。

Bayesian Elicitation with LLMs: Model Size Helps, Extra "Reasoning" Doesn't Always

  • 用不同规模模型估算统计量并给出可信区间
  • 95%区间实际覆盖率仅9-44%,严重高估信心
  • 用保形预测可校正过自信,适合做决策前处理

大型语言模型(LLMs)被提出作为人类专家的替代,用于估计未知量及其不确定性,这一过程称为贝叶斯引出。我们测试了十一款LLM在估计健康患病率、人格特质分布和劳动力市场数据等人口统计指标时的表现,并要求它们以95%可信区间表达不确定性。通过调整模型的推理强度(低、中、高),检验更多“思考”是否提升结果。结果显示:第一,更大、更强大的模型能提供更准确的估计,但增加推理努力并未带来稳定改善;第二,所有模型均严重高估信心:其95%可信区间实际包含真实值的比例仅为9%至44%,远低于预期的95%;第三,一种名为保形预测(conformal prediction)的统计校准技术可有效纠正过自信问题,通过扩大区间实现目标覆盖率。初步实验发现,给予模型网络搜索访问能力会降低已有较准确模型的表现,而对较弱模型有轻微改善。模型在常见话题上表现良好,但在专业医疗数据上表现不佳。这些结果表明,未经校准的LLM不确定性估计不能直接用于决策。

原文摘要 · Abstract (English)

Large language models (LLMs) have been proposed as alternatives to human experts for estimating unknown quantities with associated uncertainty, a process known as Bayesian elicitation. We test this by asking eleven LLMs to estimate population statistics, such as health prevalence rates, personality trait distributions, and labor market figures, and to express their uncertainty as 95\% credible intervals. We vary each model's reasoning effort (low, medium, high) to test whether more "thinking" improves results. Our findings reveal three key results. First, larger, more capable models produce more accurate estimates, but increasing reasoning effort provides no consistent benefit. Second, all models are severely overconfident: their 95\% intervals contain the true value only 9--44\% of the time, far below the expected 95\%. Third, a statistical recalibration technique called conformal prediction can correct this overconfidence, expanding the intervals to achieve the intended coverage. In a preliminary experiment, giving models web search access degraded predictions for already-accurate models, while modestly improving predictions for weaker ones. Models performed well on commonly discussed topics but struggled with specialized health data. These results indicate that LLM uncertainty estimates require statistical correction before they can be used in decision-making.

贝叶斯推断不确定性量化大模型评估保形预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。