提出可解释的类别表征,让扩散模型当老师教学生更鲁棒。
Canonical Latent Representations in Conditional Diffusion Models
- 从条件扩散模型中提取保留关键类别的紧凑潜在码
- 仅用10%训练数据大小的表征,实现强抗攻击与泛化能力
- 适合需要可解释性和鲁棒性的下游学习任务
条件扩散模型(CDMs)在生成任务中表现优异,其对完整数据分布的建模能力为下游判别学习提供了分析-合成的新路径。然而,这种建模能力也导致类别特征与无关上下文纠缠,难以提取鲁棒且可解释的表示。为此,我们提出标准潜在表征(CLAReps),即内部特征保留核心类别信息、剔除非判别信号的潜在代码。解码后,CLAReps能生成每类代表性样本,以最小无关细节呈现核心类别语义。基于此,我们构建了新型扩散蒸馏范式CaDistill:学生模型虽可访问全部训练集,但教师(CDM)仅通过CLAReps传递核心类别知识,其体积仅为原始数据的10%。训练后,学生模型展现出更强的对抗鲁棒性与泛化能力,更聚焦于类别信号而非虚假背景线索。结果表明,CDMs不仅能生成图像,还可作为紧凑、可解释的教师,驱动鲁棒表示学习。
原文摘要 · Abstract (English)
Conditional diffusion models (CDMs) have shown impressive performance across a range of generative tasks. Their ability to model the full data distribution has opened new avenues for analysis-by-synthesis in downstream discriminative learning. However, this same modeling capacity causes CDMs to entangle the class-defining features with irrelevant context, posing challenges to extracting robust and interpretable representations. To this end, we identify Canonical LAtent Representations (CLAReps), latent codes whose internal CDM features preserve essential categorical information while discarding non-discriminative signals. When decoded, CLAReps produce representative samples for each class, offering an interpretable and compact summary of the core class semantics with minimal irrelevant details. Exploiting CLAReps, we develop a novel diffusion-based feature-distillation paradigm, CaDistill. While the student has full access to the training set, the CDM as teacher transfers core class knowledge only via CLAReps, which amounts to merely 10 % of the training data in size. After training, the student achieves strong adversarial robustness and generalization ability, focusing more on the class signals instead of spurious background cues. Our findings suggest that CDMs can serve not just as image generators but also as compact, interpretable teachers that can drive robust representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。