数据集大小显著影响隐私文本重写效果,越大越平衡隐私与效用。
With Privacy, Size Matters: On the Importance of Dataset Size in Differentially Private Text Rewriting
- 通过动态分割大规模数据集,测试不同规模下隐私保护效果。
- 在百万级文本数据上发现:数据越多,隐私-效用权衡越优。
- 呼吁更严格的评估标准,对实际应用有重要指导意义。
近期在差分隐私与自然语言处理(DP NLP)领域的研究提出了多种有前景的文本重写机制。然而,这些机制的评估中常忽视一个关键因素——数据集规模对机制效用和隐私保护能力的影响。本文首次将数据集规模纳入DP文本私有化评估框架,设计了基于大规模数据集的效用与隐私测试,使用最多一百万条文本进行实验,重点量化数据规模增加对隐私-效用权衡的影响。结果表明,数据集规模在评估DP文本重写机制中起着核心作用;研究还呼吁提升DP NLP评估的严谨性,并为该领域在实际应用中的规模化发展提供新思路。
原文摘要 · Abstract (English)
Recent work in Differential Privacy with Natural Language Processing (DP NLP) has proposed numerous promising techniques in the form of text rewriting mechanisms. In the evaluation of these mechanisms, an often-ignored aspect is that of dataset size, or rather, the effect of dataset size on a mechanism's efficacy for utility and privacy preservation. In this work, we are the first to introduce this factor in the evaluation of DP text privatization, where we design utility and privacy tests on large-scale datasets with dynamic split sizes. We run these tests on datasets of varying size with up to one million texts, and we focus on quantifying the effect of increasing dataset size on the privacy-utility trade-off. Our findings reveal that dataset size plays an integral part in evaluating DP text rewriting mechanisms; additionally, these findings call for more rigorous evaluation procedures in DP NLP, as well as shed light on the future of DP NLP in practice and at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。