预训练模型通过后验收缩实现通用先验,解决经验贝叶斯问题
Universal priors: solving empirical Bayes via Bayesian inference and pretraining
- 基于后验收缩理论,发现可泛化的通用先验存在
- 在泊松经验贝叶斯问题上达到近最优后悔界˜O(1/n)
- 解释模型长度外推能力,适合贝叶斯建模与预训练研究者
我们从理论上解释了[Te et al., 2025]的实证发现:在合成数据上预训练的Transformer模型在经验贝叶斯(EB)任务中表现优异。不分析模型结构或训练动态,而是探讨为何在特定训练分布下训练的贝叶斯估计器能适应任意测试分布。聚焦泊松经验贝叶斯问题,我们证明存在通用先验,使得在此类先验下训练的估计器对所有测试分布均实现近最优后悔界˜O(1/n)。分析依赖于贝叶斯统计中的后验收缩经典现象,表明预训练Transformer通过后验收缩自适应未知测试分布。该视角也解释了长度泛化现象——测试序列长度超过训练长度,因模型执行使用广义后验的贝叶斯推断。
原文摘要 · Abstract (English)
We theoretically justify the recent empirical finding of [Teh et al., 2025] that a transformer pretrained on synthetically generated data achieves strong performance on empirical Bayes (EB) problems. We take an indirect approach to this question: rather than analyzing the model architecture or training dynamics, we ask why a pretrained Bayes estimator, trained under a prespecified training distribution, can adapt to arbitrary test distributions. Focusing on Poisson EB problems, we identify the existence of universal priors such that training under these priors yields a near-optimal regret bound of $\widetilde{O}(\frac{1}{n})$ uniformly over all test distributions. Our analysis leverages the classical phenomenon of posterior contraction in Bayesian statistics, showing that the pretrained transformer adapts to unknown test distributions precisely through posterior contraction. This perspective also explains the phenomenon of length generalization, in which the test sequence length exceeds the training length, as the model performs Bayesian inference using a generalized posterior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。