arXiv:2507.05147cond-mat.stat-mechcond-mat.dis-nn2025-07被引 4

伪似然训练让神经网络具备泛化能力的联想记忆功能

Pseudo-likelihood produces associative memories able to generalize, even for asymmetric couplings

  • 用伪似然替代传统似然,解决归一化难题
  • 小样本下模式成吸引子,吸引域比经典霍普菲尔德更大
  • 适用于多种数据集,实现从记忆到泛化的跨越

基于能量的概率模型在最大化数据似然时受限于难以计算的分区函数。常用方法是最大化伪似然,将全局归一化替换为可计算的局部归一化。本文发现,在零温度极限下,通过最大化伪似然训练的网络自然实现了联想记忆:当训练集较小时,模式成为固定点吸引子,其吸引域超过任何经典霍普菲尔德规则。我们对无关联随机模式定量解释了这一现象。此外,对于来自计算机科学(随机特征模型、MNIST)、物理学(自旋玻璃)和生物学(蛋白质)的不同结构化数据集,随着训练样本增加,学习到的网络超越单纯记忆,形成与测试样本具有非平凡相关性的有意义吸引子,展现出泛化能力。结果表明,伪似然不仅是一种高效的推断工具,更是一种原理性的记忆与泛化机制。

原文摘要 · Abstract (English)

Energy-based probabilistic models learned by maximizing the likelihood of the data are limited by the intractability of the partition function. A widely used workaround is to maximize the pseudo-likelihood, which replaces the global normalization with tractable local normalizations. Here we show that, in the zero-temperature limit, a network trained to maximize pseudo-likelihood naturally implements an associative memory: if the training set is small, patterns become fixed-point attractors whose basins of attraction exceed those of any classical Hopfield rule. We explain quantitatively this effect on uncorrelated random patterns. Moreover, we show that, for different structured datasets coming from computer science (random feature model, MNIST), physics (spin glasses) and biology (proteins), as the number of training examples increases the learned network goes beyond memorization, developing meaningful attractors with non-trivial correlations with test examples, thus showing the ability to generalize. Our results therefore reveal pseudo-likelihood works both as an efficient inference tool and as a principled mechanism for memory and generalization.

联想记忆伪似然泛化能力能量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。