arXiv:2507.06056cs.CLcs.AI2025-07中稿 · TMLR被引 1

用数据压缩性量化大模型记忆行为,发现熵与记忆呈线性关系。

Data Compressibility Quantifies LLM Memorization

  • 从实例级转为集合级压缩度量,提升可解释性
  • 数据集熵与记忆分数呈线性相关,验证了EM线性规律
  • 为理解模型记忆机制提供可量化的新工具

大型语言模型(LLMs)会记忆训练数据的片段,甚至在适当提示下复现原文。尽管关注度高,现有研究对训练数据如何影响记忆缺乏深入理解,且缺少量化刻画。本文基于数据压缩性量化记忆的研究路径,分析了以往方法失效的原因,发现将衡量尺度从实例级转向集合级后,能揭示一个稳健现象——熵-记忆(EM)线性关系:集合级数据熵估计值与记忆分数呈现线性相关。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are known to memorize portions of their training data, sometimes even reproduce content verbatim when prompted appropriately. Despite substantial interest, existing LLM memorization research has offered limited insight into how training data influences memorization and largely lacks quantitative characterization. In this work, we build upon the line of research that seeks to quantify memorization through data compressibility. We analyze why prior attempts fail to yield a reliable quantitative measure and show that a surprisingly simple shift from instance-level to set-level metrics uncovers a robust phenomenon, which we term the \textit{Entropy--Memorization (EM) Linearity}. This law states that a set-level data entropy estimator exhibits a linear correlation with memorization scores.

大模型记忆机制数据压缩熵分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。