Transformer可实现贝叶斯预测中的分布学习,突破传统点预测限制。
Transformers Can Learn Posterior Predictive Distributions In-Context

- 通过构造性证明,Transformer能模拟梯度下降求解后验预测均值与方差
- 注意力深度与分箱分辨率影响近似误差,且在超样本范围外仍可泛化
- 归一化和注意力层数选择对模型外推能力至关重要,适合研究概率建模者
近期提出的先验数据拟合网络(PFNs)在贝叶斯预测任务中表现优异,通过上下文学习逼近后验预测分布(PPD)。尽管其经验性能出色且超越点预测,但对Transformer在上下文中学习分布的算法能力尚缺乏理论理解。本文聚焦高斯过程回归问题,通过构造性证明显示,Transformer可实现针对后验预测均值与方差的梯度下降算法,并经非线性映射生成PPD的分箱概率。我们分析了近似PPD的误差界,涉及注意力深度与分箱分辨率。结果表明,归一化及注意力深度的选择在超越预训练样本规模范围的外推能力中起关键作用。仿真验证了上述发现,揭示了PFNs在目标PPD时的表达能力,以及架构选择如何影响泛化性能。
原文摘要 · Abstract (English)
Prior-data fitted networks (PFNs) have recently emerged as a powerful approach for Bayesian prediction tasks, approximating the posterior predictive distribution (PPD) through in-context learning. Despite their strong empirical performance and ability to go beyond point predictions, theoretical understandings of the algorithmic capability of transformers to learn distributions in context are still lacking. Focusing on Gaussian process regression problems, we show by construction that transformers can implement a gradient descent algorithm targeting the posterior predictive mean and variance, followed by nonlinear mappings that yield binned probabilities of PPD. We study the error bounds of the approximated PPD in terms of attention depth and bin resolution. Based on these results, we further demonstrate the key role of normalization and the choice of attention depth in enabling the extrapolation abilities of transformers beyond the pretraining sample size range. We conduct simulations that corroborate our findings, providing insight into the expressivity of PFNs targeting PPDs and how architectural choices may influence generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。