提出SOFT方法,用选择性模糊数据保护微调大模型隐私
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
- 通过选择性模糊高影响数据,动态平衡隐私与性能
- 在六领域多模型上显著降低成员推断攻击成功率
- 适合需保护训练数据隐私的AI应用开发者
大型语言模型在众多应用中取得显著成功,但微调过程常涉及敏感信息,引发严重隐私问题。本文首次全面评估微调后大模型对成员推断攻击(MIAs)的脆弱性,实证发现:攻击利用微调中的损失下降,能高效揭示成员信息。为此,我们提出SOFT(Selective Data Obfuscation in LLM Fine-Tuning),一种新防御机制,通过可调参数选择性模糊高影响力数据,在保持模型性能的同时缓解隐私泄露。实验覆盖六个不同领域及多种大模型架构与规模,结果表明SOFT能有效降低隐私风险,同时维持竞争力,提供一种实用且可扩展的微调数据保护方案。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable success and are widely adopted for diverse applications. However, fine-tuning these models often involves private or sensitive information, raising critical privacy concerns. In this work, we conduct the first comprehensive study evaluating the vulnerability of fine-tuned LLMs to membership inference attacks (MIAs). Our empirical analysis demonstrates that MIAs exploit the loss reduction during fine-tuning, making them highly effective in revealing membership information. These findings motivate the development of our defense. We propose SOFT (\textbf{S}elective data \textbf{O}bfuscation in LLM \textbf{F}ine-\textbf{T}uning), a novel defense technique that mitigates privacy leakage by leveraging influential data selection with an adjustable parameter to balance utility preservation and privacy protection. Our extensive experiments span six diverse domains and multiple LLM architectures and scales. Results show that SOFT effectively reduces privacy risks while maintaining competitive model performance, offering a practical and scalable solution to safeguard sensitive information in fine-tuned LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。