用差分隐私生成合成文本数据,解决医疗等高敏感领域数据共享难题
Evaluating Differentially Private Synthetic Data Generation in High-Stakes Domains
- 基于差分隐私语言模型生成真实高敏感领域合成数据
- 发现以往评估方法未能揭示合成数据的效用与公平性缺陷
- 为医疗、社工等领域的隐私保护数据共享提供新思路
文本数据匿名化困难严重制约了医疗、社会服务等高敏感领域自然语言处理技术的发展与应用。未经妥善匿名的数据难以共享给标注人员或外部研究者,也无法用于训练公开模型。本文探索使用差分隐私语言模型生成的合成数据替代真实数据,以在不泄露隐私的前提下推动该领域NLP发展。不同于以往研究,我们针对真实高敏感领域生成合成数据,并提出并实施了面向实际应用的评估方法来衡量数据质量。结果表明,先前简单的评估方式未能暴露合成数据在实用性、隐私性和公平性上的问题。整体而言,本工作强调需进一步改进合成数据生成技术,才能实现真正可行的隐私保护数据共享。
原文摘要 · Abstract (English)
The difficulty of anonymizing text data hinders the development and deployment of NLP in high-stakes domains that involve private data, such as healthcare and social services. Poorly anonymized sensitive data cannot be easily shared with annotators or external researchers, nor can it be used to train public models. In this work, we explore the feasibility of using synthetic data generated from differentially private language models in place of real data to facilitate the development of NLP in these domains without compromising privacy. In contrast to prior work, we generate synthetic data for real high-stakes domains, and we propose and conduct use-inspired evaluations to assess data quality. Our results show that prior simplistic evaluations have failed to highlight utility, privacy, and fairness issues in the synthetic data. Overall, our work underscores the need for further improvements to synthetic data generation for it to be a viable way to enable privacy-preserving data sharing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。