arXiv:2512.03238cs.CRcs.AI2025-12被引 14

用差分隐私生成真实用户数据的合成版本,既保护隐私又可用。

How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

  • 通过差分隐私技术生成保留原数据趋势的合成数据
  • 支持图像、表格、文本等多模态数据生成,提供强隐私保障
  • 适合需要高隐私安全的数据研究与产品开发团队

高质量数据是释放AI对终端用户潜力的关键。然而,获取新数据源日益困难:多数公开的人类生成数据已被广泛使用。此外,公开数据往往不能代表特定系统的真实用户行为——例如,研究人员使用的语音数据集来自合同工与AI助手的互动,其语言更规范、同质化程度更高,远不如真实用户的指令多样。因此,如何基于真实用户交互生成高质量数据成为关键挑战。直接使用用户数据存在重大隐私风险。差分隐私(DP)是一种成熟的隐私保护框架,可有效限制信息泄露,是保护用户隐私的行业标准。本文聚焦于“差分隐私合成数据”,即在保留原始数据整体趋势的同时,为贡献者提供强隐私保障的合成数据。这类数据能解锁因隐私顾虑而无法使用的数据集价值,并替代过去仅依赖规则化匿名化的敏感数据。本文系统梳理了差分隐私合成数据的全套技术,涵盖各类模态(图像、表格、文本、去中心化)的隐私保护能力与最新进展。文章详述了构建此类系统所需全部组件,包括敏感数据处理、数据预处理、隐私预算追踪及实证隐私测试。期望推动该技术的广泛应用,促进更多研究,并增强对差分隐私合成数据方法的信任。

原文摘要 · Abstract (English)

High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated data will soon have been used. Additionally, publicly available data often is not representative of users of a particular system -- for example, a research speech dataset of contractors interacting with an AI assistant will likely be more homogeneous, well articulated and self-censored than real world commands that end users will issue. Therefore unlocking high-quality data grounded in real user interactions is of vital interest. However, the direct use of user data comes with significant privacy risks. Differential Privacy (DP) is a well established framework for reasoning about and limiting information leakage, and is a gold standard for protecting user privacy. The focus of this work, \emph{Differentially Private Synthetic data}, refers to synthetic data that preserves the overall trends of source data,, while providing strong privacy guarantees to individuals that contributed to the source dataset. DP synthetic data can unlock the value of datasets that have previously been inaccessible due to privacy concerns and can replace the use of sensitive datasets that previously have only had rudimentary protections like ad-hoc rule-based anonymization. In this paper we explore the full suite of techniques surrounding DP synthetic data, the types of privacy protections they offer and the state-of-the-art for various modalities (image, tabular, text and decentralized). We outline all the components needed in a system that generates DP synthetic data, from sensitive data handling and preparation, to tracking the use and empirical privacy testing. We hope that work will result in increased adoption of DP synthetic data, spur additional research and increase trust in DP synthetic data approaches.

差分隐私合成数据隐私保护数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。