arXiv:2508.07054cs.LGcs.AI2025-08EMNLP被引 4

研究大模型蒸馏中的隐私泄露问题,发现学生模型会继承教师的隐私风险。

Membership and Memorization in LLM Knowledge Distillation

  • 系统分析六种大模型蒸馏方法的隐私风险来源
  • 所有蒸馏方法均存在成员推理与记忆化隐私风险
  • 不同模块、任务和训练数据导致风险差异显著

近期知识蒸馏(KD)旨在通过将大型语言模型(LLM)的知识从大“教师”模型迁移到小“学生”模型,以缓解大模型高计算成本的问题。然而,当教师模型在私有数据上训练时,学生模型可能继承其隐私风险。本文系统地刻画并研究了六种LLM KD技术中存在的成员身份和记忆化隐私风险。在涵盖七个NLP任务的指令微调设置下,结合GPT-2、LLAMA-2和OPT三类教师模型及多种尺寸的学生模型,我们证明所有现有LLM KD方法均会将教师的隐私风险传递给学生。但不同蒸馏方法的风险程度各异。我们系统分析了关键蒸馏组件(如KD目标函数、学生训练数据和任务类型)对隐私风险的影响,并揭示了记忆化与成员身份风险之间存在显著不一致性。此外,我们还对每个模型块的隐私风险进行了表征,发现不同模块间的隐私风险差异极大。

原文摘要 · Abstract (English)

Recent advances in Knowledge Distillation (KD) aim to mitigate the high computational demands of Large Language Models (LLMs) by transferring knowledge from a large ''teacher'' to a smaller ''student'' model. However, students may inherit the teacher's privacy when the teacher is trained on private data. In this work, we systematically characterize and investigate membership and memorization privacy risks inherent in six LLM KD techniques. Using instruction-tuning settings that span seven NLP tasks, together with three teacher model families (GPT-2, LLAMA-2, and OPT), and various size student models, we demonstrate that all existing LLM KD approaches carry membership and memorization privacy risks from the teacher to its students. However, the extent of privacy risks varies across different KD techniques. We systematically analyse how key LLM KD components (KD objective functions, student training data and NLP tasks) impact such privacy risks. We also demonstrate a significant disagreement between memorization and membership privacy risks of LLM KD techniques. Finally, we characterize per-block privacy risk and demonstrate that the privacy risk varies across different blocks by a large margin.

隐私安全知识蒸馏大模型成员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。