用风险最小化构建联合先验,实现更可靠的跨数据集统计推断。
Formal Bayesian Transfer Learning via the Total Risk Prior
- 以源数据条件下的风险最小化构造联合先验,统一处理源与目标数据
- 在基因数据上表现优于经典频数方法,尤其当源数据有限时
- 支持贝叶斯不确定性量化和模型平均,适合小样本跨领域研究
在数据严重受限的情况下,利用同领域中辅助的源数据集信息可显著提升统计推断效果。然而,现有迁移学习方法难以应对源数据本身有限且与目标数据不匹配的情形。传统做法是将源数据的经验损失最小化结果作为目标参数的先验均值,但这使源参数估计脱离贝叶斯框架。本文核心贡献是采用条件于源参数的风险最小化来构建先验,从而建立一个涵盖所有源数据集与目标数据集的联合先验分布。由此可实现完整的贝叶斯不确定性量化,并通过吉布斯采样对每个源数据集是否纳入模型进行平均。我们证明该先验的一种具体形式等价于变换坐标系下的贝叶斯岭回归;同时提出计算技巧,使方法可扩展至中等规模数据集。此外,我们指出近期提出的极小极大频数迁移学习方法可视为本模型的近似最大后验解。在基因组学应用中,该方法在源数据有限时表现出更优的预测性能,显著优于基准频数方法。
原文摘要 · Abstract (English)
In analyses with severe data-limitations, augmenting the target dataset with information from ancillary datasets in the application domain, called source datasets, can lead to significantly improved statistical procedures. However, existing methods for this transfer learning struggle to deal with situations where the source datasets are also limited and not guaranteed to be well-aligned with the target dataset. A typical strategy is to use the empirical loss minimizer on the source data as a prior mean for the target parameters, which places the estimation of source parameters outside of the Bayesian formalism. Our key conceptual contribution is to use a risk minimizer conditional on source parameters instead. This allows us to construct a single joint prior distribution for all parameters from the source datasets as well as the target dataset. As a consequence, we benefit from full Bayesian uncertainty quantification and can perform model averaging via Gibbs sampling over indicator variables governing the inclusion of each source dataset. We show how a particular instantiation of our prior leads to a Bayesian Lasso in a transformed coordinate system and discuss computational techniques to scale our approach to moderately sized datasets. We also demonstrate that recently proposed minimax-frequentist transfer learning techniques may be viewed as an approximate Maximum a Posteriori approach to our model. Finally, we demonstrate superior predictive performance relative to the frequentist baseline on a genetics application, especially when the source data are limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。