发现大模型记忆学术引用呈分层结构,引用越多越准确,但年份难记。
Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- 用引用次数衡量数据冗余,分析模型记忆层级
- 引用超90次后准确率跃升,超1200次近乎复述原文
- 标题作者最先记住,年份几乎不学,标题相似会混淆
大型语言模型(LLMs)在多种任务中生成流畅文本,但虚构不存在的学术引用是其严重且广为人知的缺陷。本文基于先前研究将幻觉与逐字记忆视为同一概率过程的结果,以引用次数作为训练数据冗余的代理指标,探讨单一文献记录内部的冗余结构。使用GPT-4.1生成并人工验证了跨20个计算机科学领域的100条引用,通过与真实元数据的余弦相似度衡量事实准确性。结果表明:(i) 事实准确性在不同领域间差异显著,且随引用次数呈对数线性增长;(ii) 模型在约90次引用处出现拐点,在接近1200次引用时达到饱和,此后记录几乎完全复现;(iii) 记忆具有层级性,标题和第一作者最早被回忆,而会议名称和数值字段需更高冗余,发表年份几乎无法学习;(iv) 即使高引用文献也可能因标题与作者重叠而混淆,体现为虚假吸引子干扰。因此,大模型的记忆并非简单开关状态,而是由预训练语料中知识分布不均所塑造的渐进、分层现象。
原文摘要 · Abstract (English)
Large language models (LLMs) generate fluent text across a wide range of tasks, but the fabrication of non-existent academic citations remains a critical and well-documented failure mode. Building on prior work that frames hallucination and verbatim memorization as outcomes of the same probabilistic process, this study uses citation count as a proxy for training data redundancy and asks how this redundancy is internally structured within a single bibliographic record. Using GPT-4.1, we generated and manually verified 100 citations across twenty computer-science domains, measuring factual fidelity via cosine similarity against authentic metadata. We find that (i) factual accuracy varies substantially across domains and scales log-linearly with citation count, (ii) the model crosses two empirically identifiable thresholds; an inflection around 90 citations and a saturation point near 1,200 citations beyond which records are reproduced nearly verbatim, (iii) memorization is hierarchical, with titles and first authors recalled earliest while venues and numeric fields require far greater redundancy and publication years remain essentially unlearned, and (iv) even highly cited records can be conflated when their titles and authors overlap, an effect interpretable as spurious-attractor interference. Memorization in LLMs is therefore not a binary on/off state but a graduated, hierarchically layered phenomenon shaped by the uneven distribution of knowledge in the pretraining corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。