arXiv:2512.00638cs.LG2025-12被引 2

保护隐私的扩散模型,生成混合类型表格数据

Privacy Preserving Diffusion Models for Mixed-Type Tabular Data Generation

  • 用嵌入表示类别特征,降低编码开销
  • 在同等隐私下,性能比基线高16%-42%
  • 适合金融医疗等敏感数据共享场景

我们提出DP-FinDiff,一种用于合成混合类型表格数据的差分隐私扩散框架。该方法采用基于嵌入的类别特征表示,减少编码开销,可扩展至高维数据集。为适配扩散过程中的差分隐私训练,我们提出两种隐私感知训练策略:一种自适应时间步采样器,使更新与扩散动态对齐;一种特征聚合损失,减轻截断带来的偏差。两项改进共同提升了生成数据保真度与下游任务效用,同时不削弱隐私保证。在金融和医疗数据集上,DP-FinDiff在相近隐私水平下,相比差分隐私基线实现了16%-42%的性能提升,展现出在敏感领域安全高效数据共享的巨大潜力。

原文摘要 · Abstract (English)

We introduce DP-FinDiff, a differentially private diffusion framework for synthesizing mixed-type tabular data. DP-FinDiff employs embedding-based representations for categorical features, reducing encoding overhead and scaling to high-dimensional datasets. To adapt DP-training to the diffusion process, we propose two privacy-aware training strategies: an adaptive timestep sampler that aligns updates with diffusion dynamics, and a feature-aggregated loss that mitigates clipping-induced bias. Together, these enhancements improve fidelity and downstream utility without weakening privacy guarantees. On financial and medical datasets, DP-FinDiff achieves 16-42% higher utility than DP baselines at comparable privacy levels, demonstrating its promise for safe and effective data sharing in sensitive domains.

差分隐私扩散模型数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。