arXiv:2604.07486cs.CRcs.AI2026-04ACL被引 2

用私有种子生成高保真且隐私安全的合成文本数据

Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation

论文配图:Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
图 1 · 摘自论文原文
  • 用私有种子结合差分隐私机制生成合成数据
  • 在保持数据真实性的前提下实现强隐私保护
  • 适合需要合规生成敏感文本的应用场景

大型语言模型(LLMs)已成为合成数据生成的强大工具。一个关键应用场景是生成私有文本的合成副本,需在隐私与数据效用之间取得平衡。本文提出真实且隐私保护的合成数据生成方法(RPSG),利用私有种子并集成隐私保护策略,包括在候选选择阶段引入形式化差分隐私(DP)机制,以生成高保真的合成数据。与当前最先进的私有合成数据生成方法相比,全面实验表明,RPSG 在保持对私有数据高保真度的同时,提供了强大的隐私保护。

原文摘要 · Abstract (English)

Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and utility. We propose Realistic and Privacy-Preserving Synthetic Data Generation (RPSG), which uses private seeds and integrates privacy-preserving strategies, including a formal differential privacy (DP) mechanism in the candidate selection, to generate realistic synthetic data. Comprehensive experiments against state-of-the-art private synthetic data generation methods demonstrate that RPSG achieves high fidelity to private data while providing strong privacy protection.

合成数据差分隐私LLM隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。