arXiv:2501.19158cs.LGcond-mat.dis-nn2025-01ICML被引 7

揭示能量模型过拟合根源,提供早停与修正策略。

A theoretical framework for overfitting in energy-based modeling

  • 基于耦合矩阵谱分解,分析学习时序与数据有限性的关系。
  • 发现早停最优点由学习时序与初始条件共同决定。
  • 提出适用于离散模型的实证收缩修正法,可推广至通用能量模型。

我们研究了数据有限性对用于识别相互作用网络的成对能量基模型训练的影响。以高斯模型为测试基准,通过耦合矩阵的特征基分解训练轨迹,利用特征模态独立演化特性,揭示学习时序与经验协方差矩阵的谱分解密切相关。结果表明,早停的最优时机源于这些时序与训练初始条件的相互作用。此外,我们证明有限样本修正可通过渐近随机矩阵理论准确建模,并在能量模型框架下给出广义交叉验证的对应方法。该解析框架可微调扩展至二元变量最大熵成对模型。这些发现为离散变量模型提供通过经验收缩校正控制过拟合的策略,改善能量基生成模型中的过拟合管理。最后,通过推导得分匹配算法下得分函数的神经正切核动力学,将方法推广至任意能量基模型。

原文摘要 · Abstract (English)

We investigate the impact of limited data on training pairwise energy-based models for inverse problems aimed at identifying interaction networks. Utilizing the Gaussian model as testbed, we dissect training trajectories across the eigenbasis of the coupling matrix, exploiting the independent evolution of eigenmodes and revealing that the learning timescales are tied to the spectral decomposition of the empirical covariance matrix. We see that optimal points for early stopping arise from the interplay between these timescales and the initial conditions of training. Moreover, we show that finite data corrections can be accurately modeled through asymptotic random matrix theory calculations and provide the counterpart of generalized cross-validation in the energy based model context. Our analytical framework extends to binary-variable maximum-entropy pairwise models with minimal variations. These findings offer strategies to control overfitting in discrete-variable models through empirical shrinkage corrections, improving the management of overfitting in energy-based generative models. Finally, we propose a generalization to arbitrary energy-based models by deriving the neural tangent kernel dynamics of the score function under the score-matching algorithm.

能量模型过拟合统计物理机器学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。