提出新下界,让基于JSD的表示学习更可靠。
Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning
- 用JSD推导出KLD的新紧致下界,连接两者关系。
- 最大化JSD能保证提升互信息下界,理论更扎实。
- 适合做表示学习、信息瓶颈等任务的研究者参考。
互信息(MI)是表征学习中衡量统计依赖性的基础指标。尽管直接通过其作为Kullback-Leibler散度(KLD)的定义优化MI通常不可行,许多近期方法转而最大化替代依赖度量,尤其是通过判别损失最大化联合分布与边缘乘积之间的Jensen-Shannon散度(JSD)。然而,这些代理目标与MI之间的联系仍不清晰。本文首次在一般情况下推导出一个新、紧致且可计算的KLD关于JSD的下界。将其应用于联合与边缘分布时,证明最大化JSD可提升互信息的保证下界。此外,我们重新审视JSD目标的实际实现,发现最小化区分联合对与边缘对的二分类器交叉熵,等价于已知的JSD变分下界。大量实验表明,该下界在互信息估计中具有紧致性。与现有神经估计器对比,我们的下界估计器始终提供稳定、低方差的紧致下界。我们在信息瓶颈框架中也验证了其实际效用。结果为基于判别学习的互信息表征学习提供了新的理论支持和强实证依据。
原文摘要 · Abstract (English)
Mutual Information (MI) is a fundamental measure of statistical dependence widely used in representation learning. While direct optimization of MI via its definition as a Kullback-Leibler divergence (KLD) is often intractable, many recent methods have instead maximized alternative dependence measures, most notably, the Jensen-Shannon divergence (JSD) between joint and product of marginal distributions via discriminative losses. However, the connection between these surrogate objectives and MI remains poorly understood. In this work, we bridge this gap by deriving a new, tight, and tractable lower bound on KLD as a function of JSD in the general case. By specializing this bound to joint and marginal distributions, we demonstrate that maximizing the JSD-based information increases a guaranteed lower bound on mutual information. Furthermore, we revisit the practical implementation of JSD-based objectives and observe that minimizing the cross-entropy loss of a binary classifier trained to distinguish joint from marginal pairs recovers a known variational lower bound on the JSD. Extensive experiments demonstrate that our lower bound is tight when applied to MI estimation. We compared our lower bound to state-of-the-art neural estimators of variational lower bound across a range of established reference scenarios. Our lower bound estimator consistently provides a stable, low-variance estimate of a tight lower bound on MI. We also demonstrate its practical usefulness in the context of the Information Bottleneck framework. Taken together, our results provide new theoretical justifications and strong empirical evidence for using discriminative learning in MI-based representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。