通过激活调控降低大模型记忆风险,兼顾性能与安全
Mitigating Memorization in LLMs using Activation Steering
- 直接干预模型激活值,抑制训练数据记忆
- 在Gemma上有效压制文学内容记忆,性能损失小
- 适合关注隐私保护与模型安全的研究者
大语言模型对训练数据的过度记忆会带来隐私泄露和版权内容复现等风险。激活调控是一种直接干预模型内部激活值的技术,具有操纵模型行为的潜力。本文在受控的文学材料记忆基准上评估该方法,证明其能在Gemma模型上有效抑制记忆内容,同时保持良好泛化能力。研究还分析了压制效果与语言流畅度之间的权衡,揭示了基于激活干预的优缺点。成果为构建更安全、更隐私友好的大模型提供了实用高效的解决方案。
原文摘要 · Abstract (English)
The memorization of training data by Large Language Models (LLMs) poses significant risks, including privacy leaks and the regurgitation of copyrighted content. Activation steering, a technique that directly intervenes in model activations, has emerged as a promising approach for manipulating LLMs. In this work, we explore the effectiveness of activation steering in reducing memorization while preserving generalization capabilities. We conduct empirical evaluations using a controlled memorization benchmark of literary material and demonstrate that our method successfully suppresses memorized content with minimal degradation in model performance in Gemma. Additionally, we analyze the trade-offs between suppression effectiveness and linguistic fluency, highlighting the advantages and limitations of activation-based interventions. Our findings contribute to ongoing efforts in developing safer and more privacy-preserving LLMs by providing a practical and efficient mechanism to mitigate unintended memorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。