arXiv:2410.00751cs.CL2024-10EMNLP被引 11

挑战隐私保护范式,揭示差分隐私在文本处理中的局限性

Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting

  • 用大模型提示词重写文本实现隐私保护
  • 实验证明差分隐私会显著降低文本可用性
  • 适合关注隐私技术边界的研究者参考

随着大语言模型的普及,隐私保护自然语言处理受到越来越多关注。近期文献普遍采用差分隐私(DP)集成到NLP方法中。本文对这类方法提出批判性审视,分析其带来的限制及相应挑战。以最新提出的DP-Prompt方法为例,该方法利用语言模型重写文本实现文本隐私化。我们在此基础上,在有无差分隐私条件下对比多种场景下的重写效果。通过实证评估文本效用与隐私性,结果表明:当前差分隐私在自然语言处理中的应用存在显著可用性损失,亟需更深入讨论其在实际场景中的有效性与优势。

原文摘要 · Abstract (English)

The field of privacy-preserving Natural Language Processing has risen in popularity, particularly at a time when concerns about privacy grow with the proliferation of Large Language Models. One solution consistently appearing in recent literature has been the integration of Differential Privacy (DP) into NLP techniques. In this paper, we take these approaches into critical view, discussing the restrictions that DP integration imposes, as well as bring to light the challenges that such restrictions entail. To accomplish this, we focus on $\textbf{DP-Prompt}$, a recent method for text privatization leveraging language models to rewrite texts. In particular, we explore this rewriting task in multiple scenarios, both with DP and without DP. To drive the discussion on the merits of DP in NLP, we conduct empirical utility and privacy experiments. Our results demonstrate the need for more discussion on the usability of DP in NLP and its benefits over non-DP approaches.

差分隐私文本生成大模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。