提出新型自由能函数,揭示正反KL差异的方差与尾部敏感机制。
Surprisal-Rényi Free Energy
- 基于似然比的对数矩构造新自由能,跳出传统f散度框架。
- 在正反KL极限间建立局部展开,首阶修正项为对数似然比方差。
- 兼具大偏差控制能力与最小描述长度解释,适合理论学习研究者。
前向与反向Kullback-Leibler(KL)散度在学习与推断中作为极限目标出现,但引发截然不同的归纳偏置,仅靠期望层面无法解释。本文引入一种名为惊喜-Rényi自由能(SRFE)的新函数,其为似然比的对数矩函数,不属于f-散度类。我们证明SRFE可退化为前向与反向KL散度的端点极限,并在两个极限附近导出局部展开,其中对数似然比的方差作为一阶修正项出现。这揭示了偏离KL主导范式时的显式均值-方差权衡。进一步地,我们建立了SRFE的Gibbs型变分表征:它是加权KL散度和的唯一极小化器;并证明SRFE可直接通过Chernoff型界控制超额码长的大偏差,从而给出精确的最小描述长度解释。综合这些结果,SRFE被确认为一种对方差与尾部敏感的自由能泛函,澄清了前向与反向KL极限背后的几何与大偏差结构,而无需统一或涵盖不同学习框架。
原文摘要 · Abstract (English)
The forward and reverse Kullback-Leibler (KL) divergences arise as limiting objectives in learning and inference yet induce markedly different inductive biases that cannot be explained at the level of expectations alone. In this work, we introduce the Surprisal-Rényi Free Energy (SRFE), a log-moment-based functional of the likelihood ratio that lies outside the class of $f$-divergences. We show that SRFE recovers forward and reverse KL divergences as singular endpoint limits and derive local expansions around both limits in which the variance of the log-likelihood ratio appears as a first-order correction. This reveals an explicit mean-variance tradeoff governing departures from KL-dominated regimes. We further establish a Gibbs-type variational characterization of SRFE as the unique minimizer of a weighted sum of KL divergences and prove that SRFE directly controls large deviations of excess code-length via Chernoff-type bounds, yielding a precise Minimum Description Length interpretation. Together, these results identify SRFE as a variance- and tail-sensitive free-energy functional that clarifies the geometric and large-deviation structure underlying forward and reverse KL limits, without unifying or subsuming distinct learning frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。