提出概率电路曲率的组合理论,揭示全局正则化缺陷并设计更优的局部正则方法。
A Compositional Theory of Curvature in Probabilistic Circuits

- 将曲率分解为节点使用度与局部尖锐度乘积,实现精确可计算的组合分析
- 发现全局正则化会因深度偏差导致欠拟合,而新方法提升泛化能力
- 支持闭式更新,适合需要精确推理的概率电路建模任务
概率电路(PCs)是支持精确推理的生成模型,与深度神经网络不同,其损失曲面曲率可通过对数似然的海森迹精确且高效地衡量。近期工作通过全局正则化该迹,以引导学习趋向平坦、泛化性更好的极值点。我们证明,将尖锐度视为全局正则项在概率电路中可能误设,因其曲率具有内在的组合特性。我们严格证明,每个求和节点对海森迹的贡献可精确分解为其电路流(衡量节点使用频率)与由输出分布决定的局部尖锐度项的乘积。这一分解揭示了全局尖锐度正则化存在深度偏差,并可能导致欠拟合。基于此,我们提出一种自适应尖锐度感知正则器,根据节点固有局部曲率施加惩罚,同时保持闭式期望最大化(EM)更新。实验表明,该方法在保留尖锐度感知学习的鲁棒性与优势的同时,恢复了全局正则化所牺牲的泛化性能。
原文摘要 · Abstract (English)
Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. Recent work regularizes this trace globally to bias learning toward flatter, better generalizing optima. We show that treating sharpness as a global regularizer can be misspecified for PCs, whose curvature is inherently compositional. We prove that each sum node's contribution to the Hessian trace factorizes exactly into its circuit flow, which measures how heavily the node is used, and a local sharpness term determined by its output distribution. This decomposition provides insights into why global sharpness regularization is depth biased and can lead to underfitting. Building on it, we introduce an adaptive sharpness aware regularizer that penalizes nodes based on intrinsic local curvature and preserves closed form EM updates. We also show that empirically, this targeted regularization recovers the generalization that global regularization sacrifices while retaining the robustness and benefits of sharpness aware learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。