微调大模型易泄露隐私数据,新方法可零泄漏保95%性能
Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
- 通过四类互补技术减少数据记忆,防止敏感信息泄露
- 重复敏感数据导致泄露率升至60%-75%,较基线提升64.2%
- 适合关注模型隐私安全的研究者与企业应用开发者
大型语言模型在自然语言处理任务中表现卓越,但其在微调过程中容易记忆训练数据,带来严重隐私风险。本文对微调后的大型语言模型中的数据记忆现象进行了全面实证分析,并提出一种多层隐私保护框架。在GPT-2、Phi-3和Gemma-2等现代模型上开展控制实验,结果表明:使用重复敏感数据进行微调,使隐私泄露率从基线的0%-5%上升至60%-75%,跨测试模型平均提升64.2%。我们提出并严格评估了四种互补的隐私保护方法:语义数据去重、生成阶段差分隐私、基于熵的过滤以及基于模式的内容过滤。实验显示,这些方法可将数据泄露降至0%,同时保留原模型94.7%的性能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks, but their tendency to memorize training data poses significant privacy risks, particularly during fine-tuning processes. This paper presents a comprehensive empirical analysis of data memorization in fine-tuned LLMs and introduces a novel multi-layered privacy protection framework. Through controlled experiments on modern LLM architectures including GPT-2, Phi-3, and Gemma-2, we demonstrate that fine-tuning with repeated sensitive data increases privacy leakage rates from baseline levels of 0-5% to 60-75%, representing a 64.2% average increase across tested models. We propose and rigorously evaluate four complementary privacy protection methods: semantic data deduplication, differential privacy during generation, entropy-based filtering, and pattern-based content filtering. Our experimental results show that these techniques can reduce data leakage to 0% while maintaining 94.7% of original model utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。