揭示大模型记忆与提取的隐私盲区,证明现有审计方法可能失效。
Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots

- 提出精确的差分隐私边界,区分记忆与可提取性。
- 实验证明百万参数模型仍存在审计盲区,伪造提示可复现完整内容。
- 为隐私审计提供新标准,适合关注模型安全的研究者。
大型语言模型中的记忆现象缺乏统一定义,差分隐私(DP)常被用作统一保护手段。本文精确刻画了两类实际重要的度量——反事实记忆和自适应提取——的差分隐私常数,并证明二者互不控制。在 f-DP 框架下,任意自适应提取协议在列表预算为 m 时,成功概率不超过 1−f(κ),其中 κ 为无差别基线;该界在密集基线上紧致。最小熵可无前提地保证:当纯 ε- DP 下,对任意先验分布,只要风险水平 τ≤1/2,均有 H∞≥ε log₂e + log₂(m/τ)。对于记忆侧,f-DP 将任意有界得分的记忆优势上限控制在函数 η(f) 内,纯 DP 下等于 tanh(ε/2);当存在 ≥2 个重复副本时,简单的 ε↦kε 上界 tanh(kε/2) 不可达,真实常数为几何噪声计数构造的闭式阶梯函数。该上界可在实践中使用的局部得分类中达到,正是在此处,两种度量分离:某些机制虽被记忆但不可提取,另一些则完全可提取却对所有基于损失的评分完全不可见。这种双向盲区在十亿参数模型中依然存在:通过保留触发器的释放可从单个提示中原样恢复内容,而当前实践中的审计工具却仍将其标记为‘干净’。
原文摘要 · Abstract (English)
Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is treated as a proxy against all of them at once. We pin down the exact DP constant for the two that carry the practical weight, counterfactual memorization and adaptive extraction, and show that they do not control each other. Under $f$-DP, every adaptive extraction protocol with list budget $m$ succeeds with probability at most $1-f(κ)$ for the oblivious baseline $κ$, and the bound is tight on a dense set of baselines: DP uniformly controls extraction exactly up to a threshold in how well the secret can be guessed a priori. Min-entropy certifies that baseline distribution-free, since $H_\infty\geε\log_2 e+\log_2(m/τ)$ holds extraction below a risk level $τ\le1/2$ under pure $ε$-DP for every prior, and is exact on uniform priors. On the memorization side, $f$-DP caps the counterfactual memorization of any bounded score at an advantage functional $η(f)$, equal to $\tanh(ε/2)$ under pure DP; for $k\ge2$ duplicated copies the naive $ε\mapsto kε$ bound $\tanh(kε/2)$ is unattainable, the exact constant being a closed-form staircase attained by geometric noisy counting. That cap is attained inside the local score class used in practice, and it is there that the two measures separate: one mechanism is memorized yet unextractable, another fully extractable yet exactly invisible to every loss-based score. The two-sided blind spot this opens for loss-based auditing and unlearning verification survives on billion-parameter models: a reserved-trigger release is recovered verbatim from one prompt while the audits practitioners deploy certify it clean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。