用修正后的AI生成数据构建先验,提升小样本下的推断准确性
Supercharging Bayesian Inference with Reliable AI-Informed Priors

- 通过修正AI生成数据的偏差来构建更可靠的先验
- 相比传统方法显著降低后验偏倚,提高可信区间覆盖率
- 适用于小样本场景,尤其适合医疗等高可靠性需求领域
现代预测系统所蕴含的信念可在数据有限时作为统计推断的先验信息。利用预测模型构建信息丰富的先验虽能增强小样本推断效果,但可能将模型误差带入后验分布。本文提出一种AI驱动的先验构造框架,通过修正生成合成数据的AI规律,避免误差传播。该修正后的规律可嵌入基于合成数据的先验构造方法,如作为狄利克雷过程先验中数据生成过程的基测度。我们称由此产生的先验及其后验为修正AI先验与修正AI后验,并在非消失先验强度下建立了其高斯渐近性质,推导出中心化偏差的一阶表达式。实验表明,修正后的先验显著降低偏倚、改善可信区间覆盖率,使AI提供的先验信息更可靠。此外,在真实皮肤疾病分类任务中应用该先验,显著提升了预测性能。
原文摘要 · Abstract (English)
Modern predictive systems encode beliefs that can act as useful prior information for statistical inference in data-limited settings. Using them for prior construction introduces a tradeoff: an informative prior built from a predictive model can sharpen inference from limited data, but also risks propagating error from the model into the posterior. We propose a framework for AI-informed prior elicitation that mitigates this tension by rectifying the AI-induced law that generates synthetic data before using it to inform a prior. The rectified law can be embedded into synthetic data-driven prior elicitation techniques, including as a base measure in a Dirichlet process (DP) prior on the data-generating process. We refer to the resulting prior and corresponding posterior as the rectified AI prior and rectified AI posterior. We establish Gaussian asymptotics for the rectified AI posterior under non-vanishing prior strength and derive a first-order expression for its centering bias. Our rectified AI priors substantially reduce bias compared to standard approaches, improve the coverage of credible intervals, and make AI-powered prior information more reliable. We additionally apply the rectified AI prior to a real skin disease classification task and show that it can meaningfully boost predictive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。