梳理大模型微调中的隐私风险与防护方法,揭示数据泄露隐患。
Privacy in Fine-tuning Large Language Models: Attacks, Defenses, and Future Directions
- 系统分析微调阶段的三类隐私攻击机制
- 评估差分隐私等防御手段在保持性能下的有效性
- 适合关注模型安全与合规的开发者与研究者
微调已成为利用大语言模型(LLMs)完成特定下游任务的关键步骤,使模型在多个领域达到顶尖性能。然而,微调过程常涉及敏感数据,暴露于多种隐私风险之中,这些风险利用了该阶段的独特特性。本文全面综述了微调过程中与隐私相关的挑战,重点分析了成员推理、数据提取和后门攻击等隐私攻击形式。同时,回顾了针对微调阶段隐私风险的防御机制,包括差分隐私、联邦学习和知识遗忘等方法,讨论其在缓解隐私风险与维持模型效用方面的效果与局限性。通过识别现有研究的关键空白,本文指出了当前挑战,并提出了推动隐私保护微调方法发展的未来方向,以促进大模型在多样化应用中的负责任使用。
原文摘要 · Abstract (English)
Fine-tuning has emerged as a critical process in leveraging Large Language Models (LLMs) for specific downstream tasks, enabling these models to achieve state-of-the-art performance across various domains. However, the fine-tuning process often involves sensitive datasets, introducing privacy risks that exploit the unique characteristics of this stage. In this paper, we provide a comprehensive survey of privacy challenges associated with fine-tuning LLMs, highlighting vulnerabilities to various privacy attacks, including membership inference, data extraction, and backdoor attacks. We further review defense mechanisms designed to mitigate privacy risks in the fine-tuning phase, such as differential privacy, federated learning, and knowledge unlearning, discussing their effectiveness and limitations in addressing privacy risks and maintaining model utility. By identifying key gaps in existing research, we highlight challenges and propose directions to advance the development of privacy-preserving methods for fine-tuning LLMs, promoting their responsible use in diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。