arXiv:2606.10481cs.LGcs.AI2026-06

用合成数据检测大模型隐私泄露,提升审计精度与可复现性。

Advancing the State-of-the-Art in Empirical Privacy Auditing

论文配图:Advancing the State-of-the-Art in Empirical Privacy Auditing
图 1 · 摘自论文原文
  • 通过高温采样生成高影响力合成探针,增强隐私泄露检测能力。
  • 在真实数据集上实现90%以上成员推理攻击成功率的精确量化。
  • 适用于需验证训练数据隐私性的模型开发者和合规审查人员。

大型语言模型(LLM)的参数高效微调可能对个别训练样本产生过度记忆。经验隐私审计(EPA)通过测量成员推断(MI)或重构攻击中的实际数据泄露来量化这一风险。关键挑战在于设计与敏感训练数据混合的“探针”示例。本文提出通过从LLM中使用高温度采样(T ≥ 0.8)生成合成探针,结合针对敏感数据定制的提示词。这些探针作为高影响力异常点,确保高可识别性,从而实现强审计效果。此外,由于探针本身不具隐私性,可重复插入且不影响真实数据隐私。一个重要应用场景是基于敏感数据微调的模型生成合成数据,但同样存在隐私风险。本文引入一种基于辅助模型微调的合成数据审计方法:将辅助模型在合成数据上微调后,检测其对原始探针的响应,从而提供合成数据泄露的强估计。最后,利用所提强审计方法,系统研究了模型容量与探针熵对记忆效应的交互影响。

原文摘要 · Abstract (English)

Parameter-efficient fine-tuning of large language models (LLMs) can exhibit problematic memorization of individual training examples. Empirical privacy auditing (EPA) quantifies this risk by measuring realistic data leakage on membership inference (MI) or reconstruction attacks. A key challenge in EPA is designing ``canary'' examples that are mixed with the privacy-sensitive training data. We propose generating synthetic canaries via high-temperature sampling ($T \geq 0.8$) from LLMs, using prompts tailored to the privacy-sensitive training data. These canaries act as high-influence outliers, ensuring high identifiability and hence strong audits. Further, since the canaries are themselves non-private, they are inspectable and can be inserted with repetition without jeopardizing the privacy of the real data. An important use of models fine-tuned on privacy-sensitive data is the generation of synthetic data. This also comes with privacy risk. We introduce a powerful synthetic data audit based on fine-tuning an auxiliary model on the synthetic data. Auditing the auxiliary model for the original canaries then provides a strong estimate of the privacy leakage through the synthetic data. Finally, leveraging our strong auditing methodologies, we perform a systematic investigation into the interacting effects of model capacity and canary entropy on memorization.

隐私审计大模型合成数据成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。