arXiv:2508.01749cs.CVcs.AI2025-08ICCV被引 5

提升隐私保护数据蒸馏中的噪声效率,用更少数据达到更好效果

Improving Noise Efficiency in Privacy-preserving Dataset Distillation

  • 分离采样与优化过程,改善收敛性
  • 在CIFAR-10上用50张图/类提升10.0%
  • 适合关注隐私保护与数据压缩的研究者

现代机器学习模型依赖大规模数据集,其中常包含敏感信息,引发严重隐私问题。差分隐私(DP)数据生成通过创建合成数据集,在预设隐私预算下限制私密信息泄露;但其需要大量数据才能达到原始数据训练模型的性能。为降低合成数据生成成本,数据蒸馏(DD)因其极高的训练与存储效率脱颖而出。该效率在结合DP机制时尤为有利,可在不牺牲隐私的前提下生成紧凑且信息丰富的合成数据集。然而,现有最先进的隐私化数据蒸馏方法存在采样与优化同步进行、依赖随机初始化网络带来的噪声信号的问题,导致私密信息利用效率低下,因添加过多噪声而浪费。为此,我们提出一种新框架,将采样与优化解耦以促进更好收敛,并通过在信息子空间中匹配来降低DP噪声影响,提升信号质量。在CIFAR-10上,本方法在每类50张图像时实现10.0%的性能提升,在仅需前人方法五分之一蒸馏集大小时仍取得8.3%的增益,展现出显著推动隐私保护数据蒸馏的潜力。

原文摘要 · Abstract (English)

Modern machine learning models heavily rely on large datasets that often include sensitive and private information, raising serious privacy concerns. Differentially private (DP) data generation offers a solution by creating synthetic datasets that limit the leakage of private information within a predefined privacy budget; however, it requires a substantial amount of data to achieve performance comparable to models trained on the original data. To mitigate the significant expense incurred with synthetic data generation, Dataset Distillation (DD) stands out for its remarkable training and storage efficiency. This efficiency is particularly advantageous when integrated with DP mechanisms, curating compact yet informative synthetic datasets without compromising privacy. However, current state-of-the-art private DD methods suffer from a synchronized sampling-optimization process and the dependency on noisy training signals from randomly initialized networks. This results in the inefficient utilization of private information due to the addition of excessive noise. To address these issues, we introduce a novel framework that decouples sampling from optimization for better convergence and improves signal quality by mitigating the impact of DP noise through matching in an informative subspace. On CIFAR-10, our method achieves a \textbf{10.0\%} improvement with 50 images per class and \textbf{8.3\%} increase with just \textbf{one-fifth} the distilled set size of previous state-of-the-art methods, demonstrating significant potential to advance privacy-preserving DD.

隐私保护数据蒸馏差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。