智能分配隐私预算,让文本改写更安全且保留语义
Spend Your Budget Wisely: Towards an Intelligent Distribution of the Privacy Budget in Differentially Private Text Rewriting
- 根据语言特征为文本不同部分动态分配隐私预算
- 相同预算下,智能分配可提升隐私保护效果
- 适合关注隐私安全与文本质量平衡的研究者
差分隐私文本重写任务旨在对敏感文本进行重写,在保证差分隐私(DP)的前提下隐藏显性和隐性身份标识,同时保持原文语义。近年来,该领域涌现出多种词级、句级和文档级的DP重写方法。这些方法普遍依赖一个隐私预算参数(即ε),决定文本的隐私化程度。然而,由于语言结构的特殊性,现有方法忽视了隐私预算应如何在文本中合理分配——并非所有语言成分都具有同等敏感性。本文首次提出针对这一问题的解决方案,构建并评估了一套基于语言学和NLP的方法,用于将隐私预算智能分配至文本中的各个词元。通过一系列隐私与效用实验,我们证明:在相同隐私预算条件下,智能分配相比均匀分配能实现更高的隐私保护水平,并带来更优的隐私-效用权衡。本研究揭示了文本差分隐私保护的复杂性,呼吁进一步探索高效利用DP隐私优势的方法。
原文摘要 · Abstract (English)
The task of $\textit{Differentially Private Text Rewriting}$ is a class of text privatization techniques in which (sensitive) input textual documents are $\textit{rewritten}$ under Differential Privacy (DP) guarantees. The motivation behind such methods is to hide both explicit and implicit identifiers that could be contained in text, while still retaining the semantic meaning of the original text, thus preserving utility. Recent years have seen an uptick in research output in this field, offering a diverse array of word-, sentence-, and document-level DP rewriting methods. Common to these methods is the selection of a privacy budget (i.e., the $\varepsilon$ parameter), which governs the degree to which a text is privatized. One major limitation of previous works, stemming directly from the unique structure of language itself, is the lack of consideration of $\textit{where}$ the privacy budget should be allocated, as not all aspects of language, and therefore text, are equally sensitive or personal. In this work, we are the first to address this shortcoming, asking the question of how a given privacy budget can be intelligently and sensibly distributed amongst a target document. We construct and evaluate a toolkit of linguistics- and NLP-based methods used to allocate a privacy budget to constituent tokens in a text document. In a series of privacy and utility experiments, we empirically demonstrate that given the same privacy budget, intelligent distribution leads to higher privacy levels and more positive trade-offs than a naive distribution of $\varepsilon$. Our work highlights the intricacies of text privatization with DP, and furthermore, it calls for further work on finding more efficient ways to maximize the privatization benefits offered by DP in text rewriting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。