提出快速稀疏贝叶斯学习的普适剪枝准则,揭示其内在机制。
General Pruning Criteria for Fast SBL
- 通过分析单个超参数对边缘似然的影响,推导出剪枝条件。
- 在高斯假设下,条件与Fast SBL剪枝规则完全一致。
- 为稀疏贝叶斯学习的剪枝行为提供理论解释,适合模型压缩研究者。
稀疏贝叶斯学习(SBL)为线性模型中每个权重分配一个超参数,假设权重服从均值为零、精度(方差倒数)等于对应超参数的高斯分布。该方法通过积分消去权重并进行边缘最大似然估计来求解超参数。传统SBL常导致多个超参数估计趋于无穷,使对应权重被置零(即剪枝),从而获得权重向量的稀疏估计。本文在减弱噪声和权重分布的高斯假设前提下,分析单个超参数变化时的边缘似然函数,推导出导致超参数有限或无穷的充分条件。结果表明,在高斯情形下,这两个条件互补且等价于快速SBL(F-SBL)的剪枝条件,从而为该算法提供了新的理论视角。
原文摘要 · Abstract (English)
Sparse Bayesian learning (SBL) associates to each weight in the underlying linear model a hyperparameter by assuming that each weight is Gaussian distributed with zero mean and precision (inverse variance) equal to its associated hyperparameter. The method estimates the hyperparameters by marginalizing out the weights and performing (marginalized) maximum likelihood (ML) estimation. SBL returns many hyperparameter estimates to diverge to infinity, effectively setting the estimates of the corresponding weights to zero (i.e., pruning the corresponding weights from the model) and thereby yielding a sparse estimate of the weight vector. In this letter, we analyze the marginal likelihood as function of a single hyperparameter while keeping the others fixed, when the Gaussian assumptions on the noise samples and the weight distribution that underlies the derivation of SBL are weakened. We derive sufficient conditions that lead, on the one hand, to finite hyperparameter estimates and, on the other, to infinite ones. Finally, we show that in the Gaussian case, the two conditions are complementary and coincide with the pruning condition of fast SBL (F-SBL), thereby providing additional insights into this algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。