arXiv:2510.07452cs.CRcs.CL2025-10Conference of the …被引 3

精准定位并修改语言模型中的隐私泄露电路,提升隐私保护效果。

PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing

  • 通过电路发现技术定位导致隐私泄露的具体计算路径
  • 可将隐私信息泄露召回率降低最高达65%
  • 可与差分隐私结合,实现接近0的残留泄露

语言模型可能从训练数据中记忆个人身份信息(PII),在推理阶段被攻击者提取。现有防御方法如差分隐私虽能降低泄露,但严重损害模型性能。基于对计算电路的全面分析,我们发现特定的PII泄露电路是根源所在。为此提出PATCH(隐私感知的定向电路修补):先识别再直接编辑这些电路以减少泄露。实验表明,PATCH在隐私-效用权衡上优于现有方法,可使模型的PII泄露召回率降低高达65%。此外,与差分隐私结合后,残余泄露可降至0.01%。分析显示,现有防御措施无法消除这些泄露电路,而PATCH能有效抑制其影响。

原文摘要 · Abstract (English)

Language models (LMs) may memorize personally identifiable information (PII) from training data, enabling adversaries to extract it during inference. Existing defense mechanisms such as differential privacy (DP) reduce this leakage, but incur large drops in utility. Based on a comprehensive study using circuit discovery to identify the computational circuits responsible PII leakage in LMs, we hypothesize that specific PII leakage circuits in LMs should be responsible for this behavior. Therefore, we propose PATCH (Privacy-Aware Targeted Circuit PatcHing), a novel approach that first identifies and subsequently directly edits PII circuits to reduce leakage. PATCH achieves better privacy-utility trade-off than existing defenses, e.g., reducing recall of PII leakage from LMs by up to 65%. Finally, PATCH can be combined with DP to reduce recall of residual leakage of an LM to as low as 0.01%. Our analysis shows that PII leakage circuits persist even after the application of existing defense mechanisms. In contrast, PATCH can effectively mitigate their impact.

隐私保护语言模型电路修复差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。