arXiv:2605.10795stat.MLcond-mat.dis-nn2026-05

揭示线性记忆网络存储事实的极限与最优机制。

Factual recall in linear associative memories: sharp asymptotics and mechanistic insights

论文配图:Factual recall in linear associative memories: sharp asymptotics and mechanistic insights
图 1 · 摘自论文原文
  • 构建解耦模型分析关联存储约束,发现容量受输入输出分离度限制。
  • 理论证明可存储约 $p_c \log p_c / d^2 = 1/2$ 对关联,与维度相关。
  • 揭示最优学习规则通过提升正确得分超越赫布法则,具可解释性。

大型语言模型展现出惊人的事实回忆能力,但神经网络存储与检索输入-输出关联的根本极限仍不明确。本文研究一个极简设置:一个线性关联记忆网络,将 $p$ 个 $bR^d$ 中的输入嵌入映射到对应的 $d$ 维目标,要求每个映射后的输入与所有其他目标充分分离。不同于监督分类,这种严格分离带来每对关联 $p$ 个约束,导致约束间强相关,使存储容量难以直接刻画。本文引入一个解耦模型,其中每个输入拥有独立的竞争输出集,并通过数值和解析证据表明该模型在存储容量、权重谱及存储机制上与原模型等价。利用统计物理工具,证明解耦模型可存储至多 $p_c \log p_c / d^2 = 1/2$ 对关联,并推广至线性两层架构。分析还揭示最优解如何改进朴素赫布学习:不是泛化提升输入-输出对齐,而是将正确得分提升至由竞争输出决定的极端值阈值之上。这些发现为线性网络中的事实存储提供了精确的统计物理表征,为理解更现实神经架构的记忆容量提供基准。

原文摘要 · Abstract (English)

Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output associations with neural networks remain unclear. We study these limits in a minimal setting: a linear associative memory that maps $p$ input embeddings in $\mathbb{R}^d$ to their corresponding~$d$-dimensional targets via a single layer, requiring each mapped input to be well separated from all other targets. Unlike in supervised classification, this strict separation induces~$p$ constraints per association and produces strong correlations between constraints that make a direct characterisation of the storage capacity difficult. Here, we provide a precise characterisation of this capacity in the following way. We first introduce a decoupled model in which each input has its own independent set of competing outputs, and provide numerical and analytical evidence that this decoupled model is equivalent to the original model in terms of storage capacity, spectra of the learnt weights, and storage mechanism. Using tools from statistical physics, we show that the decoupled model can store up to $p_c \log p_c / d^2 = 1 / 2$ associations, and generalise the computation of $p_c$ to linear two-layer architectures. Our analysis also gives mechanistic insight into how the optimal solution improves over a naïve Hebbian learning rule: rather than boosting input-output alignments with broad fluctuations, the optimal solution raises the correct scores just above the extreme-value threshold set by the competing outputs. These findings give a sharp statistical-physics characterisation of factual storage in linear networks and provide a baseline for understanding the memory capacity of more realistic neural architectures.

记忆机制线性网络统计物理存储容量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。