arXiv:2508.20736cs.CL2025-08EMNLP被引 2

用语义三元组实现低隐私预算下的私有文档生成

Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guarantees

  • 基于语义三元组构建邻域感知的隐私生成机制
  • 在低ε值下仍保持文本连贯性,隐私与可用性平衡更好
  • 适合对隐私要求高且需保持文本质量的应用场景

许多结合差分隐私(DP)与自然语言处理的研究通过在DP保障下转换文本以保护隐私。这可通过词扰动或全文重写等方式实现,通常采用本地差分隐私(Local DP)。在本地DP下,输入文本必须在隐私参数ε的约束范围内,与任何其他潜在文本不可区分。然而,近期研究显示,本地DP下的文本隐私化仅在极高ε值时才能合理实现。为此,本文提出DP-ST方法,利用语义三元组实现邻域感知的私有文档生成,满足本地DP保证。评估表明,划分-解决范式有效,尤其当将DP概念限制在私有邻域范围时。结合大语言模型后处理,该方法可在较低ε值下生成连贯文本,同时兼顾隐私与实用性。结果强调了连贯性在合理ε水平下实现平衡隐私输出的重要性。

原文摘要 · Abstract (English)

Many works at the intersection of Differential Privacy (DP) in Natural Language Processing aim to protect privacy by transforming texts under DP guarantees. This can be performed in a variety of ways, from word perturbations to full document rewriting, and most often under local DP. Here, an input text must be made indistinguishable from any other potential text, within some bound governed by the privacy parameter $\varepsilon$. Such a guarantee is quite demanding, and recent works show that privatizing texts under local DP can only be done reasonably under very high $\varepsilon$ values. Addressing this challenge, we introduce DP-ST, which leverages semantic triples for neighborhood-aware private document generation under local DP guarantees. Through the evaluation of our method, we demonstrate the effectiveness of the divide-and-conquer paradigm, particularly when limiting the DP notion (and privacy guarantees) to that of a privatization neighborhood. When combined with LLM post-processing, our method allows for coherent text generation even at lower $\varepsilon$ values, while still balancing privacy and utility. These findings highlight the importance of coherence in achieving balanced privatization outputs at reasonable $\varepsilon$ levels.

差分隐私文本生成语义三元组本地DP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。