arXiv:2509.11176cs.CLcs.AI2025-09被引 1

差分隐私训练的文本生成模型会显著降低输出质量。

Differentially-private text generation degrades output language quality

  • 用差分隐私微调大模型,控制隐私强度
  • 更强隐私下文本长度减少77%以上,语法错误增9%以上
  • 适合关注隐私保护但需权衡生成质量的研究者

通过在大语言模型(LLMs)上施加差分隐私(DP)进行微调以保障用户隐私,近年来日益流行。然而,这种做法对生成文本的语言质量和实用性影响尚不明确。本文在三个语料库上对五种大模型进行了四种不同隐私强度下的微调,评估了生成文本的长度、语法正确性及词汇多样性。结果表明,在更强隐私约束下,生成文本长度至少减少77%,语法错误率至少上升9%,二元词组多样性至少下降10%。此外,这些合成文本在下游分类任务(如基于书评的书籍体裁识别、基于口头尸检的死因识别)中的准确率也下降,可能影响合成数据的实际可用性。

原文摘要 · Abstract (English)

Ensuring user privacy by synthesizing data from large language models (LLMs) tuned under differential privacy (DP) has become popular recently. However, the impact of DP fine-tuned LLMs on the quality of the language and the utility of the texts they produce has not been investigated. In this work, we tune five LLMs with three corpora under four levels of privacy and assess the length, the grammatical correctness, and the lexical diversity of the text outputs they produce. We also probe the utility of the synthetic outputs in downstream classification tasks such as book genre recognition based on book descriptions and cause of death recognition based on verbal autopsies. The results indicate that LLMs tuned under stronger privacy constrains produce texts that are shorter by at least 77 %, that are less grammatically correct by at least 9 %, and are less diverse by at least 10 % in bi-gram diversity. Furthermore, the accuracy they reach in downstream classification tasks decreases, which might be detrimental to the usefulness of the generated synthetic data.

差分隐私文本生成语言质量大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。