解释能量模型为何偏好简单分布,揭示学习顺序的内在机制。
Distributional simplicity bias and effective convexity in Energy Based Models
- 用有效模型视角分析能量学习动态,发现两类不动点。
- 低阶相互作用先于高阶被学习,形成层次化训练过程。
- 揭示分布简化偏差根源,适合研究生成模型优化者阅读。
能量模型是强大的生成建模框架,但其训练本质非凸,可能导致对初始化敏感、陷入劣质局部最优及梯度动态不稳定。我们通过有效模型视角对能量学习进行动力学分析,该模型可视为具有高阶相互作用的广义Ising模型或能量的傅里叶展开。在充分表达能力下,我们证明正分布上的梯度流存在两类不动点:数据一致点(精确重现目标分布)和虚假点(满足平稳性但不匹配目标分布)。在数据一致点附近,扰动要么稳定,要么中性,中性方向保持有效模型不变。此外,我们发现梯度动力学诱导出层次结构,低阶相互作用优先于高阶被学习。这为分布简化偏差提供了机制解释,并阐明为何实践中未观察到低阶不匹配的非数据一致不动点。
原文摘要 · Abstract (English)
Energy-based learning is a powerful framework for generative modelling, but its training is inherently non-convex, leading potentially to sensitivity to initialisation, poor local optima, and unstable gradient dynamics. We present a dynamical analysis of energy-based learning through the lens of the effective model, which can be interpreted as either a generalised Ising model with higher-order interactions or the Fourier expansion of the energy. Under sufficient expressivity, we show that the gradient flow induced by learning strictly positive distributions over binary variables admits two types of fixed points: data-consistent points, which exactly reproduce the target distribution, and spurious points, which satisfy stationarity without matching the target distribution. Around data-consistent points, we show that perturbations are either stable or neutral, with neutral directions leaving the effective model invariant. Finally, we show that gradient dynamics induce a hierarchy in which lower-order interactions are learned before higher-order ones. This provides a mechanistic explanation for the distributional simplicity bias and clarifies why fixed points that are not data-consistent at low orders are not observed in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。