提出基于费舍尔-杨损失的新变分学习框架,提升模型泛化能力。
Fenchel-Young Variational Learning
- 用费舍尔-杨损失构建新型变分方法,替代传统KL散度
- 新算法在多个任务上超越经典方法,支持稀疏数据和稀疏后验
- 适用于需要自适应稀疏性的建模场景,如稀疏观测系统
从变分视角看,许多统计学习准则涉及寻找在经验风险与正则化之间平衡的分布。本文通过引入基于费舍尔-杨(Fenchel-Young, FY)损失的新通用变分方法,拓展了这一视角。这些损失被视作广义的散度,可涵盖并推广经典变分学习中的Kullback-Leibler(KL)散度。所提出的FY变分学习框架包含新的概念:FY自由能、FY证据、FY证据下界和FY后验。我们推导出交替最小化和梯度反向传播算法,用于计算或下界FY证据,从而实现比以往变分形式更广泛模型的学习。这导致了经典算法的广义FY版本,例如FY期望最大化(FYEM)算法,以及潜在变量模型如FY变分自编码器(FYVAE)。实验表明,新方法在实践中具有竞争力,常优于经典方法,且具有定性新颖特征:例如,FYEM具有自适应稀疏的E步,而FYVAE可支持稀疏观测和稀疏后验。
原文摘要 · Abstract (English)
From a variational perspective, many statistical learning criteria involve seeking a distribution that balances empirical risk and regularization. In this paper, we broaden this perspective by introducing a new general class of variational methods based on Fenchel-Young (FY) losses, treated as divergences that generalize (and encompass) the familiar Kullback-Leibler divergence at the core of classical variational learning. Our proposed formulation -- FY variational learning -- includes as key ingredients new notions of FY free energy, FY evidence, FY evidence lower bound, and FY posterior. We derive alternating minimization and gradient backpropagation algorithms to compute (or lower bound) the FY evidence, which enables learning a wider class of models than previous variational formulations. This leads to generalized FY variants of classical algorithms, such as an FY expectation-maximization (FYEM) algorithm, and latent-variable models, such as an FY variational autoencoder (FYVAE). Our new methods are shown to be empirically competitive, often outperforming their classical counterparts, and most importantly, to have qualitatively novel features. For example, FYEM has an adaptively sparse E-step, while the FYVAE can support models with sparse observations and sparse posteriors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。