发现语言模型删数据后仍能被还原,提出新方法审计隐私泄露风险。
Auditing Language Model Unlearning via Information Decomposition
- 用信息分解法拆解模型记忆,识别残留知识成分。
- 实验证明删数据后仍有可线性解码的信息,易遭攻击重建。
- 设计可解释的隐私风险评分,推理时主动规避敏感输入。
我们揭示了当前语言模型去学习方法的一个关键缺陷:尽管去学习算法看似成功,但遗忘数据的信息仍可从内部表征中线性解码。为系统评估这一矛盾,我们引入基于部分信息分解(PID)的可解释、信息论框架来审计去学习效果。通过对比去学习前后模型表征,我们将与遗忘数据的互信息分解为不同成分,形式化定义了‘已遗忘’与‘残余知识’。分析显示,冗余信息(在两模型间共享)构成残余知识,且与已知对抗重建攻击的脆弱性相关。基于此,我们提出一种基于表征的风险评分,在推理时可指导模型对敏感输入选择不响应,从而有效缓解隐私泄露。本工作提供了原理清晰、可操作的去学习表征级审计方法,为语言模型更安全部署提供理论支持与实用工具。
原文摘要 · Abstract (English)
We expose a critical limitation in current approaches to machine unlearning in language models: despite the apparent success of unlearning algorithms, information about the forgotten data remains linearly decodable from internal representations. To systematically assess this discrepancy, we introduce an interpretable, information-theoretic framework for auditing unlearning using Partial Information Decomposition (PID). By comparing model representations before and after unlearning, we decompose the mutual information with the forgotten data into distinct components, formalizing the notions of unlearned and residual knowledge. Our analysis reveals that redundant information, shared across both models, constitutes residual knowledge that persists post-unlearning and correlates with susceptibility to known adversarial reconstruction attacks. Leveraging these insights, we propose a representation-based risk score that can guide abstention on sensitive inputs at inference time, providing a practical mechanism to mitigate privacy leakage. Our work introduces a principled, representation-level audit for unlearning, offering theoretical insight and actionable tools for safer deployment of language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。