对比扩散模型与LLM在差分隐私下的文本生成能力
Private Synthetic Text Generation with Diffusion Models
- 用扩散模型生成私有化文本数据,对比其与自回归LLM表现
- 开源LLM在隐私保护下优于扩散模型,性能更稳定
- 揭露先前研究中隐私保证被违反的潜在假设
扩散模型生成合成文本的能力如何?近期研究表明其性能可媲美自回归大语言模型。但若训练过程采用差分隐私,其表现如何?目前尚无实证,然而私有图像生成的成功前景令人期待。本文通过大量实验回答这一开放问题。同时,我们重新评估并复现了以往关于私有文本生成的LLM研究,揭示其中存在未满足的假设,可能导致差分隐私保证失效。结果部分反驳了非私有场景下的结论,表明在隐私约束下,完全开源的LLM表现优于扩散模型。本文所有代码、数据集及实验配置均已公开,以促进后续研究。
原文摘要 · Abstract (English)
How capable are diffusion models of generating synthetics texts? Recent research shows their strengths, with performance reaching that of auto-regressive LLMs. But are they also good in generating synthetic data if the training was under differential privacy? Here the evidence is missing, yet the promises from private image generation look strong. In this paper we address this open question by extensive experiments. At the same time, we critically assess (and reimplement) previous works on synthetic private text generation with LLMs and reveal some unmet assumptions that might have led to violating the differential privacy guarantees. Our results partly contradict previous non-private findings and show that fully open-source LLMs outperform diffusion models in the privacy regime. Our complete source codes, datasets, and experimental setup is publicly available to foster future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。