arXiv:2502.01187cs.AIcs.CL2025-02被引 4

揭示大模型记忆的极端不均衡现象,提出量化与分解方法

Skewed Memorization in Large Language Models: Quantification and Decomposition

  • 分析微调过程中记忆概率随序列长度的变化规律
  • 发现记忆行为呈现显著偏态分布,与训练时长和数据相似性相关
  • 为隐私风险评估提供新指标,适合关注模型安全的研究者

大型语言模型(LLMs)中的记忆化现象带来隐私与安全风险,模型可能无意中复现敏感或受版权保护的数据。现有分析多聚焦平均情况,忽略了记忆分布的高度偏态特性。本文研究了在监督微调(SFT)中记忆化的行为,探讨其与训练时长、数据集规模及样本间相似性的关系。通过分析不同序列长度下的记忆概率,将这种偏态性与令牌生成过程联系起来,为记忆程度的估计和与已有度量的比较提供洞见。结合理论分析与实证评估,本文全面揭示了记忆行为特征,并提出检测与缓解风险的策略,助力构建更注重隐私保护的LLM。

原文摘要 · Abstract (English)

Memorization in Large Language Models (LLMs) poses privacy and security risks, as models may unintentionally reproduce sensitive or copyrighted data. Existing analyses focus on average-case scenarios, often neglecting the highly skewed distribution of memorization. This paper examines memorization in LLM supervised fine-tuning (SFT), exploring its relationships with training duration, dataset size, and inter-sample similarity. By analyzing memorization probabilities over sequence lengths, we link this skewness to the token generation process, offering insights for estimating memorization and comparing it to established metrics. Through theoretical analysis and empirical evaluation, we provide a comprehensive understanding of memorization behaviors and propose strategies to detect and mitigate risks, contributing to more privacy-preserving LLMs.

大模型安全记忆化隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。