arXiv:2503.19377cs.CVcs.LG2025-03CVPR被引 10

通过后处理方法实现生成模型的可解释性,训练快且无需大量标注。

Interpretable Generative Models through Post-hoc Concept Bottlenecks

  • 提出两种低成本后处理方法:概念瓶颈自编码器与概念控制器。
  • 在多个数据集上可解释性提升约25%,训练速度比之前快4-15倍。
  • 适用于GAN与扩散模型,适合需要可控生成的科研与应用开发者。

概念瓶颈模型(CBM)旨在构建依赖人类可理解概念进行预测的可解释模型。然而,现有基于CBM的可解释生成模型方法效率低、难扩展,需从头训练生成模型并依赖人工标注的概念监督。为此,本文提出两种新颖且低成本的后处理方法——概念瓶颈自编码器(CB-AE)与概念控制器(CC),可在无需真实数据、仅需极少量或无需概念标注的情况下实现高效可扩展训练。所提方法适用于现代生成模型家族,包括生成对抗网络与扩散模型。在CelebA、CelebA-HQ和CUB等标准数据集上,其可解释性与可控性显著优于先前工作,平均提升约25%,且训练速度提高4-15倍。大规模用户研究进一步验证了方法的可解释性与可控性。

原文摘要 · Abstract (English)

Concept bottleneck models (CBM) aim to produce inherently interpretable models that rely on human-understandable concepts for their predictions. However, existing approaches to design interpretable generative models based on CBMs are not yet efficient and scalable, as they require expensive generative model training from scratch as well as real images with labor-intensive concept supervision. To address these challenges, we present two novel and low-cost methods to build interpretable generative models through post-hoc techniques and we name our approaches: concept-bottleneck autoencoder (CB-AE) and concept controller (CC). Our proposed approaches enable efficient and scalable training without the need of real data and require only minimal to no concept supervision. Additionally, our methods generalize across modern generative model families including generative adversarial networks and diffusion models. We demonstrate the superior interpretability and steerability of our methods on numerous standard datasets like CelebA, CelebA-HQ, and CUB with large improvements (average ~25%) over the prior work, while being 4-15x faster to train. Finally, a large-scale user study is performed to validate the interpretability and steerability of our methods.

可解释生成概念瓶颈后处理可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。