梳理大模型记忆风险,揭示隐私泄露根源
Undesirable Memorization in Large Language Models: A Survey
- 从粒度、可检索性、有害性三维度分类记忆现象
- 指出训练数据重复内容易被模型复现导致隐私泄露
- 适合关注模型安全与隐私的开发者和研究者
尽管大型语言模型(LLMs)展现出卓越能力,但其伴随的风险同样不容忽视。其中,隐私与安全漏洞尤为严重,带来重大伦理与法律挑战。核心问题在于模型记忆现象——即模型倾向于存储并再现训练数据中的短语。该现象已被证实是多种针对LLM的隐私与安全攻击的根本来源。本文对现有文献进行分类,从粒度、可检索性、有害性三个维度系统梳理记忆问题。接着讨论量化记忆的指标与方法,分析其成因与影响因素。随后综述现有缓解策略,并展望未来研究方向,包括隐私与性能平衡方法,以及在对话代理、检索增强生成和扩散语言模型等具体场景中记忆现象的分析。鉴于该领域研究进展迅速,我们还维护一个持续更新的参考文献库,以反映最新成果。
原文摘要 · Abstract (English)
While recent research increasingly showcases the remarkable capabilities of Large Language Models (LLMs), it is equally crucial to examine their associated risks. Among these, privacy and security vulnerabilities are particularly concerning, posing significant ethical and legal challenges. At the heart of these vulnerabilities stands memorization, which refers to a model's tendency to store and reproduce phrases from its training data. This phenomenon has been shown to be a fundamental source to various privacy and security attacks against LLMs. In this paper, we provide a taxonomy of the literature on LLM memorization, exploring it across three dimensions: granularity, retrievability, and desirability. Next, we discuss the metrics and methods used to quantify memorization, followed by an analysis of the causes and factors that contribute to memorization phenomenon. We then explore strategies that are used so far to mitigate the undesirable aspects of this phenomenon. We conclude our survey by identifying potential research topics for the near future, including methods to balance privacy and performance, and the analysis of memorization in specific LLM contexts such as conversational agents, retrieval-augmented generation, and diffusion language models. Given the rapid research pace in this field, we also maintain a dedicated repository of the references discussed in this survey which will be regularly updated to reflect the latest developments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。