arXiv:2506.02371cs.LG2025-06被引 1

用连续优化框架替代传统迭代训练,提升扩散模型隐私保护能力

SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples

  • 将噪声数据训练重构成连续优化问题,消除人工协调步骤
  • 在线版本在多个基准上超越主流基线,性能更优且稳定
  • 适合关注数据隐私与高效训练的生成模型研究者

扩散模型虽生成效果优异,但常依赖包含敏感内容的大规模数据集,且易记忆训练数据引发隐私问题。现有方法SFBD通过在污染数据上训练并利用少量干净样本捕捉局部结构以提升收敛性,但其迭代去噪与微调循环需手动协调,实现复杂。本文将SFBD重新解释为交替投影算法,并提出连续变体SFBD flow,消除了交替步骤的需求。进一步揭示其与基于一致性约束方法的联系,并验证其实际版本Online SFBD在多个基准上持续优于强基线。

原文摘要 · Abstract (English)

Diffusion models achieve strong generative performance but often rely on large datasets that may include sensitive content. This challenge is compounded by the models' tendency to memorize training data, raising privacy concerns. SFBD (Lu et al., 2025) addresses this by training on corrupted data and using limited clean samples to capture local structure and improve convergence. However, its iterative denoising and fine-tuning loop requires manual coordination, making it burdensome to implement. We reinterpret SFBD as an alternating projection algorithm and introduce a continuous variant, SFBD flow, that removes the need for alternating steps. We further show its connection to consistency constraint-based methods, and demonstrate that its practical instantiation, Online SFBD, consistently outperforms strong baselines across benchmarks.

扩散模型隐私保护连续优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。