无需梯度优化,快速生成多样反事实样本以提升大模型鲁棒性
Visual Disentangled Diffusion Autoencoders: Scalable Counterfactual Generation for Foundation Models
- 通过解耦字典学习在可解释方向编辑嵌入,实现高效反事实生成
- 相比基线速度更快,可生成多个解耦的反事实样本,提升下游性能
- 适合需要缓解数据偏差、增强模型泛化能力的研究者使用
大模型虽具备强大的零样本能力,但易受虚假相关性与'聪明汉斯'策略影响。现有方法常依赖不可用的组标签或计算成本高的梯度对抗优化。为此,我们提出视觉解耦扩散自编码器(DiDAE),将冻结的大模型与解耦字典学习结合,实现针对大模型的高效、无梯度反事实生成。DiDAE 先在解耦字典的可解释方向上编辑大模型嵌入,再通过扩散自编码器解码,生成多个多样化、解耦的反事实样本,速度远超仅生成单一纠缠反事实的基线方法。结合反事实知识蒸馏后,DiDAE-CFKD 在缓解捷径学习方面达到当前最优性能,显著提升不平衡数据集上的下游表现。
原文摘要 · Abstract (English)
Foundation models, despite their robust zero-shot capabilities, remain vulnerable to spurious correlations and 'Clever Hans' strategies. Existing mitigation methods often rely on unavailable group labels or computationally expensive gradient-based adversarial optimization. To address these limitations, we propose Visual Disentangled Diffusion Autoencoders (DiDAE), a novel framework integrating frozen foundation models with disentangled dictionary learning for efficient, gradient-free counterfactual generation directly for the foundation model. DiDAE first edits foundation model embeddings in interpretable disentangled directions of the disentangled dictionary and then decodes them via a diffusion autoencoder. This allows the generation of multiple diverse, disentangled counterfactuals for each factual, much faster than existing baselines, which generate single entangled counterfactuals. When paired with Counterfactual Knowledge Distillation, DiDAE-CFKD achieves state-of-the-art performance in mitigating shortcut learning, improving downstream performance on unbalanced datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。