评估隐私保护生成文本在真实场景下的效果,揭示现有方法的局限性。
Evaluating Differentially Private Generation of Domain-Specific Text
- 构建统一基准,评估不同隐私条件下生成文本的实用性和保真度。
- 五大数据集测试显示,严格隐私约束下生成数据性能显著下降。
- 适合关注隐私安全与生成质量平衡的研究者和从业者参考。
生成式AI在医疗、金融等高风险领域具有变革潜力,但隐私与监管障碍限制了真实数据的使用。为此,差分隐私合成数据生成成为有前景的替代方案。本文提出一个统一基准,系统评估在正式差分隐私(DP)保障下生成文本数据集的效用与保真度。该基准解决了领域特定评估中的关键挑战,包括代表性数据选择、合理隐私预算设定,以及预训练影响和多种评估指标的考量。我们在五个领域特定数据集上评估了当前最先进的隐私保护生成方法,发现相比真实数据,在严格隐私约束下存在显著的效用与保真度下降。这些结果突显了现有方法的局限性,指出了对更先进隐私保护数据共享方法的需求,并为现实场景下的评估设立了先例。
原文摘要 · Abstract (English)
Generative AI offers transformative potential for high-stakes domains such as healthcare and finance, yet privacy and regulatory barriers hinder the use of real-world data. To address this, differentially private synthetic data generation has emerged as a promising alternative. In this work, we introduce a unified benchmark to systematically evaluate the utility and fidelity of text datasets generated under formal Differential Privacy (DP) guarantees. Our benchmark addresses key challenges in domain-specific benchmarking, including choice of representative data and realistic privacy budgets, accounting for pre-training and a variety of evaluation metrics. We assess state-of-the-art privacy-preserving generation methods across five domain-specific datasets, revealing significant utility and fidelity degradation compared to real data, especially under strict privacy constraints. These findings underscore the limitations of current approaches, outline the need for advanced privacy-preserving data sharing methods and set a precedent regarding their evaluation in realistic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。