arXiv:2411.07070cs.CLcs.AI2024-11被引 3

提出主动隐私审计框架,检测语言模型微调中的隐私泄露风险。

On Active Privacy Auditing in Supervised Fine-tuning for White-Box Language Models

  • 用改进的白盒成员推断攻击监测微调过程隐私风险
  • 在GPT-2、Llama2等模型上显著提升攻击效果
  • 为模型开发者提供可直接使用的隐私审计工具

预训练与微调已成为各类自然语言处理应用的主流技术。然而,近期研究发现,微调数据因敏感性、领域特性和可识别性,带来重大隐私风险。为此,我们提出名为Parsing的新型主动隐私审计框架,用于在语言模型(LMs)的监督微调(SFT)过程中识别并量化隐私泄露风险。该框架以改进的白盒成员推断攻击(MIAs)为核心,采用新颖的学习目标和两阶段流程,最大化暴露隐私风险。同时,我们在GPT-2、Llama2及其变体等大型语言模型上提升了MIAs的有效性。本研究旨在为语言模型微调社区提供可靠、即用的隐私审计工具,并揭示微调过程中的关键隐私隐患。实验结果表明,该框架在多种模型和任务中均具高效性,凸显了微调阶段的显著隐私问题。项目代码已公开于https://anonymous.4open.science/r/PARSING-4817/。

原文摘要 · Abstract (English)

The pretraining and fine-tuning approach has become the leading technique for various NLP applications. However, recent studies reveal that fine-tuning data, due to their sensitive nature, domain-specific characteristics, and identifiability, pose significant privacy concerns. To help develop more privacy-resilient fine-tuning models, we introduce a novel active privacy auditing framework, dubbed Parsing, designed to identify and quantify privacy leakage risks during the supervised fine-tuning (SFT) of language models (LMs). The framework leverages improved white-box membership inference attacks (MIAs) as the core technology, utilizing novel learning objectives and a two-stage pipeline to monitor the privacy of the LMs' fine-tuning process, maximizing the exposure of privacy risks. Additionally, we have improved the effectiveness of MIAs on large LMs including GPT-2, Llama2, and certain variants of them. Our research aims to provide the SFT community of LMs with a reliable, ready-to-use privacy auditing tool, and to offer valuable insights into safeguarding privacy during the fine-tuning process. Experimental results confirm the framework's efficiency across various models and tasks, emphasizing notable privacy concerns in the fine-tuning process. Project code available for https://anonymous.4open.science/r/PARSING-4817/.

隐私审计语言模型微调安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。