arXiv:2604.07238cs.LGcs.CL2026-04被引 4

隐私保护对语言识别与生成的影响远比想象中轻微

On the Price of Privacy for Language Identification and Generation

  • 在近似差分隐私下,语言任务误差率几乎不受影响
  • 纯差分隐私使误差率指数衰减速度降为原速的min{1,ε}倍
  • 结果适用于真实场景,对模型设计有直接指导意义

随着大语言模型越来越多地训练于敏感用户数据,理解语言学习中隐私保护的根本代价变得至关重要。本文首次在统计盲设下研究了差分隐私(DP)语言识别与生成问题,建立了匹配的算法与下界,精确量化了隐私成本。对于两类任务,在近似$(\varepsilon, δ)$-DP且$\varepsilon > 0$为常数时,可恢复非私有情况下的误差率:识别任务为$\exp(-r(n))$(任意$r(n) = o(n)$),生成任务为$\exp(-Ω(n))$。在纯$\varepsilon$-DP下,指数衰减速率退化为乘以$\min\{1, \varepsilon\}$因子,该结果在常数意义下紧致。值得注意的是,在合理假设下,生成任务的上界$\exp(-\min\{1,\varepsilon\} \cdot Ω(n))$与下界仅差常数因子,达到最优速率。结果表明,语言学习中的隐私代价出人意料地小:近似DP下基本不存在,纯DP下仅为指数项的$\min\{1,\varepsilon\}$倍。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly trained on sensitive user data, understanding the fundamental cost of privacy in language learning becomes essential. We initiate the study of differentially private (DP) language identification and generation in the agnostic statistical setting, establishing algorithms and matching lower bounds that precisely quantify the cost of privacy. For both tasks, approximate $(\varepsilon, δ)$-DP with constant $\varepsilon > 0$ recovers the non-private error rates: $\exp(-r(n))$ for identification (for any $r(n) = o(n)$) and $\exp(-Ω(n))$ for generation. Under pure $\varepsilon$-DP, the exponents degrade by a multiplicative factor of $\min\{1, \varepsilon\}$, which we show is tight up to constants. Notably, for generation under pure DP with mild assumptions, the upper bound $\exp(-\min\{1,\varepsilon\} \cdot Ω(n))$ matches the lower bound up to some constants, establishing an optimal rate. Our results show that the cost of privacy in language learning is surprisingly mild: absent entirely under approximate DP, and exactly a $\min\{1,\varepsilon\}$ factor in the exponent under pure DP.

差分隐私语言模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。