arXiv:2605.01065cs.CL2026-05中稿 · PrivateNLP 2026被引 1

优化文本分块与隐私预算分配,提升差分隐私文本混淆效果

A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation

论文配图:A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation
图 1 · 摘自论文原文
  • 系统测试不同文本分块与隐私预算分配方法的组合效果
  • 相同隐私预算下,不同设计导致显著不同的混淆结果
  • 为提升隐私-质量平衡提供可优化的实际路径,适合隐私保护研究者

差分隐私文本混淆的目标是通过对输入文本进行扰动,在保证差分隐私(DP)的前提下,使输出文本与原始文本在统计上难以区分。虽然词级扰动直观易行,但有意义的文本隐私化需作用于完整文档。近期研究已为隐私预算分配提供了理论基础,即如何将整体ε预算合理分配给文本的各个组成部分。本文系统评估了多种文本分解与预算分配技术在差分隐私文本混淆中的表现,考察不同文本分块方式与ε分配策略的组合效果。实验表明,这些设计选择极为关键:即使隐私预算相近,不同方法仍会导致显著差异的结果。本研究为通过优化差分隐私混淆流程来最大化实际权衡提供了可信证据。

原文摘要 · Abstract (English)

The goal of differentially private text obfuscation is to obfuscate, or "perturb", input texts with Differential Privacy (DP) guarantees, such that the private output texts are quantifiably indistinguishable from the originals. While perturbation at the word level is intuitive, meaningful text privatization happens on complete documents. Recent research has laid the groundwork for reasoning about privacy budget distribution, namely, how an overall $\varepsilon$ budget can be sensibly distributed among the component pieces of a text. We perform a systematic evaluation of multiple text decomposition and budget distribution techniques in the context of DP text obfuscation, testing how different methods for chunking texts can be combined with techniques for allocating $\varepsilon$ to these chunks. Our experiments reveal that such design choices are very important, as even with comparable privacy budgets, significantly different results can occur based on which methods are chosen. In this, we provide credible evidence of the feasibility of maximizing empirical trade-offs by optimizing DP obfuscation procedures.

差分隐私文本混淆预算分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。