用无标签数据提升扩散模型性能,无需额外标注。
Semi-Supervised Conditional Diffusion via Label Augmentation

- 给无标签数据分配统一标签,与有标签数据一起训练扩散模型。
- 理论证明:无标签数据越多,生成结果越快收敛,误差更小。
- 适合需要高效利用数据的图像、表格生成任务研究者。
条件扩散模型已成为从带标签数据中学习复杂条件分布的强大且灵活的框架。然而,在实际应用中,高质量标签的获取成本高、耗时长,导致大量无标签数据被闲置。为此,我们提出标签增强型条件扩散模型(LACD),通过为无标签样本赋予一个特定的平凡标签,并在扩展数据集上进行联合去噪得分匹配,从而有效融合无标签数据。我们给出了该方法在总体层面可识别目标条件分布的充分条件。此外,我们建立了严格的统计保证:当存在足够多的无标签样本时,LACD生成的采样分布在总变差距离下收敛速度严格快于纯监督估计器,且在一阶沃尔德斯坦距离下至少同样快。在合成数据、图像和表格基准上的大量实验验证了理论,并表明相较于纯监督估计器,LACD在样本效率和生成性能上均有显著提升。
原文摘要 · Abstract (English)
Conditional diffusion models have become a powerful and flexible framework for learning complex conditional distributions from labeled data. In practice, however, acquiring high-quality labels is costly and time-consuming, leaving large volumes of unlabeled data unused. To address this, we introduce label-augmented conditional diffusion (LACD), a simple and effective approach that incorporates unlabeled examples by assigning them a designated trivial label and performing joint denoising score matching over the augmented dataset. We provide sufficient conditions guaranteeing population-level identifiability of the target conditional distribution under this scheme. Moreover, we establish rigorous statistical guarantees: when sufficiently many unlabeled samples are available, the sampling distribution produced by LACD converges strictly faster than the purely supervised estimator in total variation distance, and at least as fast in Wasserstein-1 distance. Extensive experiments on synthetic, image, and tabular benchmarks corroborate our theory and show substantial gains in sample efficiency and generative performance compared with the purely supervised estimator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。