arXiv:2501.19022cs.CL2025-01NAACL被引 8

研究噪声对差分隐私文本重写的影响,发现噪声严重损害效果但保护隐私。

On the Impact of Noise in Differentially Private Text Rewriting

  • 提出新型句子补全隐私化方法,评估噪声作用
  • 非差分隐私方法保真度更高,但隐私保护弱于差分隐私
  • 揭示当前差分隐私技术的局限与替代方案机会

文本隐私化常依赖差分隐私(DP)提供形式化保障,其通常在数据或模型层面向文本向量添加受隐私参数ε控制的校准噪声。然而,噪声引入几乎必然导致显著性能损失,凸显了DP在NLP中的主要缺陷。本文提出一种新的句子补全隐私化技术,用于探究噪声在DP文本重写中的影响。实验表明,非差分隐私方法在保真度上表现更优,可实现可接受的隐私-效用权衡,但在实际隐私保护能力上仍不及差分隐私方法。结果突显了当前DP重写机制中噪声的关键影响,引发对DP在NLP中优劣及非差分隐私方法潜力的讨论。

原文摘要 · Abstract (English)

The field of text privatization often leverages the notion of $\textit{Differential Privacy}$ (DP) to provide formal guarantees in the rewriting or obfuscation of sensitive textual data. A common and nearly ubiquitous form of DP application necessitates the addition of calibrated noise to vector representations of text, either at the data- or model-level, which is governed by the privacy parameter $\varepsilon$. However, noise addition almost undoubtedly leads to considerable utility loss, thereby highlighting one major drawback of DP in NLP. In this work, we introduce a new sentence infilling privatization technique, and we use this method to explore the effect of noise in DP text rewriting. We empirically demonstrate that non-DP privatization techniques excel in utility preservation and can find an acceptable empirical privacy-utility trade-off, yet cannot outperform DP methods in empirical privacy protections. Our results highlight the significant impact of noise in current DP rewriting mechanisms, leading to a discussion of the merits and challenges of DP in NLP, as well as the opportunities that non-DP methods present.

差分隐私文本生成隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。