探究微调大模型时差分隐私的防护效果
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
- 对比多种微调方法下差分隐私的隐私保护能力
- 高隐私预算可显著降低隐私风险,但牺牲模型性能
- 部分微调方法因性能下降过快,不适用于差分隐私
微调大型语言模型(LLMs)已成为适应特定任务的关键策略;然而,这一过程会引入严重的隐私挑战,因为敏感训练数据可能被意外记忆并暴露。尽管差分隐私(DP)在理论上能有效防范此类泄露,其在实际大模型微调中的隐私保护效果仍不明确,尤其是在不同微调方法下的表现。本文系统研究了差分隐私在不同微调方法和隐私预算下的影响,通过数据提取攻击和成员推理攻击评估其实际隐私风险。主要发现:(1) 差分隐私会降低模型效用,但影响因微调方法而异;(2) 无差分隐私时,不同微调方法的隐私风险差异显著;(3) 即使隐私预算较高,应用差分隐私也能大幅降低隐私风险;(4) 不同微调方法在差分隐私训练下的隐私-效用权衡差异极大,某些方法因效用严重下降而不适合使用差分隐私。研究结果为隐私敏感的模型部署提供实践指导,并推动未来对微调中隐私-效用权衡优化的研究。
原文摘要 · Abstract (English)
Fine-tuning large language models (LLMs) has become an essential strategy for adapting them to specialized tasks; however, this process introduces significant privacy challenges, as sensitive training data may be inadvertently memorized and exposed. Although differential privacy (DP) offers strong theoretical guarantees against such leakage, its empirical privacy effectiveness on LLMs remains unclear, especially under different fine-tuning methods. In this paper, we systematically investigate the impact of DP across fine-tuning methods and privacy budgets, using both data extraction and membership inference attacks to assess empirical privacy risks. Our main findings are as follows: (1) Differential privacy reduces model utility, but its impact varies significantly across different fine-tuning methods. (2) Without DP, the privacy risks of models fine-tuned with different approaches differ considerably. (3) When DP is applied, even a relatively high privacy budget can substantially lower privacy risk. (4) The privacy-utility trade-off under DP training differs greatly among fine-tuning methods, with some methods being unsuitable for DP due to severe utility degradation. Our results provide practical guidance for privacy-conscious deployment of LLMs and pave the way for future research on optimizing the privacy-utility trade-off in fine-tuning methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。