arXiv:2604.22076cs.LGcs.CL2026-04

揭示大模型删私信息时的隐藏漏洞与浅层遗忘问题

PrivUn: Unveiling Latent Ripple Effects and Shallow Forgetting in Privacy Unlearning

论文配图:PrivUn: Unveiling Latent Ripple Effects and Shallow Forgetting in Privacy Unlearning
图 1 · 摘自论文原文
  • 设计三层次攻击评估框架,量化隐私删除效果
  • 发现删除信息会沿梯度关联扩散,且多数方法仅浅层清除
  • 提出梯度感知选核与多层干预策略,实现深层删除

大型语言模型在训练中常记忆私密信息,引发严重隐私风险。尽管机器删忆技术已出现,其对抗隐私攻击的真实有效性仍不明确。为此,我们提出PrivUn评估框架,通过三层次攻击场景(直接检索、上下文学习恢复、微调复现)结合遗忘评分、关联度量和遗忘深度分析,系统评估删忆鲁棒性。研究揭示当前删忆方法存在显著缺陷:1)删忆呈现梯度驱动的涟漪效应——不同于传统基于语义关系的遗忘(如知识图谱),隐私删忆沿潜在梯度关联传播;2)多数方法存在浅层遗忘,无法有效清除分布在多个深层模型层中的私密信息。为验证上述发现,我们探索两种新策略:基于梯度相似性的核心集选择,以及通过表征约束实现的多层深度干预。这些策略标志着从浅层遗忘向深层遗忘的范式转变。

原文摘要 · Abstract (English)

Large language models (LLMs) often memorize private information during training, raising serious privacy concerns. While machine unlearning has emerged as a promising solution, its true effectiveness against privacy attacks remains unclear. To address this, we propose PrivUn, a new evaluation framework that systematically assesses unlearning robustness through three-tier attack scenarios: direct retrieval, in-context learning recovery, and fine-tuning restoration; combined with quantitative analysis using forgetting scores, association metrics, and forgetting depth assessment. Our study exposes significant weaknesses in current unlearning methods, revealing two key findings: 1) unlearning exhibits gradient-driven ripple effects: unlike traditional forgetting which follows semantic relations (e.g., knowledge graphs), privacy unlearning propagates across latent gradient-based associations; and 2) most methods suffer from shallow forgetting, failing to remove private information distributed across multiple deep model layers. To validate these insights, we explore two strategies: association-aware core-set selection that leverages gradient similarity, and multi-layer deep intervention through representational constraints. These strategies represent a paradigm shift from shallow forgetting to deep forgetting.

隐私保护删忆大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。