一个简单KL恒等式推导出指数族核心理论,无需复杂证明。
Exponential families from a single KL identity
- 基于KL散度差的恒等式,用对偶变量统一推导
- 直接推出变分推断、强化学习、RLHF等关键公式
- 适合想理解指数族数学本质的研究者
指数族涵盖现代机器学习中的核心分布——softmax、高斯分布与玻尔兹曼分布——并支撑变分推断、熵正则强化学习及基于人类反馈的强化学习(RLHF)的理论。本文揭示了一个针对指数族的简洁恒等式:KL(q || p_{λ₂}) - KL(q || p_{λ₁}) 可表示为对数归一化函数 A(λ) 与矩 μ_q。令人惊讶的是,仅凭该恒等式与 KL ≥ 0(等号成立当且仅当 p = q)这一基本事实,通过直接代入与重排,即可推导出一系列经典结果:任意参考分布的广义三点恒等式、I-投影与反向I-投影的毕达哥拉斯定理、对数归一化函数的凸性、其勒让德对偶在KL意义下的识别、吉布斯变分原理,以及KL正则化奖励最大化的显式最优解,包括熵正则控制与RLHF背后的指数倾斜公式。除纯代数推导外,标准分析方法还可恢复对数归一化函数梯度公式、族内KL散度的Bregman表示,以及矩映射的满射性。本文自包含。
原文摘要 · Abstract (English)
Exponential families encompass the distributions central to modern machine learning -- softmax, Gaussians, and Boltzmann distributions -- and underlie the theory of variational inference, entropy-regularized reinforcement learning, and RLHF. We isolate a simple identity for exponential families that expresses the KL difference $\mathrm{KL}(q \| p_{λ_2}) - \mathrm{KL}(q \| p_{λ_1})$ in terms of the log-partition function $A(λ)$ and the moment $μ_q$. Remarkably, this identity together with the single fact that $\mathrm{KL} \geq 0$ (with equality iff $p = q$) suffices, by direct substitution and rearrangement, to derive a cluster of results that are classically obtained by separate, heavier arguments: a generalized three-point identity for arbitrary reference distributions, Pythagorean theorems for I-projections and reverse I-projections, convexity of the log-partition function, identification of its Legendre dual in KL terms, the Gibbs variational principle, and the explicit optimizer in KL-regularized reward maximization, including the exponential tilting formula underlying entropy-regularized control and RLHF. Beyond these purely algebraic consequences, standard analytic arguments recover the gradient formula for the log-partition function, the Bregman representation of within-family KL divergence, and the surjectivity of the moment map. The note is self-contained.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。