arXiv:2507.00258cs.CLcs.AI2025-07被引 2

对比不同微调方法,发现提示法更难泄露隐私数据。

Impact of Fine-Tuning Methods on Memorization in Large Language Models

  • 按提示词方式微调模型,比参数微调更抗隐私泄露
  • 提示法在不同模型规模下均保持低记忆风险
  • 适合关注模型隐私安全的研究者和开发者

随着预训练大语言模型能力持续提升,'预训练+微调'范式日益主流,催生了多种微调方法。然而,微调过程中的记忆化带来的隐私风险尚未得到足够重视。为此,我们对主流微调方法进行分类,并通过成员推理攻击(MIAs)评估其对记忆化的影响。结果表明,相较于参数微调,提示词微调在性能相当的前提下,对成员推理攻击的防御能力更强;且无论模型规模如何,提示词方法均能维持较低的记忆化水平。这说明参数微调更容易泄露私有信息,而提示词微调是更具隐私保护性的选择。

原文摘要 · Abstract (English)

As the capabilities of pre-trained large language models (LLMs) continue to advance, the "pre-train and fine-tune" paradigm has become increasingly mainstream, leading to the development of various fine-tuning methods. However, the privacy risks arising from memorization during fine-tuning have received relatively little attention. To address this gap, we categorize popular fine-tuning approaches and assess their impact on memorization through the lens of membership inference attacks (MIAs). Our results show that, compared to parameter-based fine-tuning, prompt-based fine-tuning achieves competitive performance while exhibiting lower vulnerability to MIAs. Furthermore, prompt-based methods maintain low memorization regardless of model scale. These findings suggest that parameter-based fine-tuning is more prone to leaking private information, whereas prompt-based fine-tuning serves as a more privacy-preserving option.

大模型隐私保护微调方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。