提出新方法让模型在增量学习中保持概念与类别的动态关系,提升可解释性与性能。
Walking the Web of Concept-Class Relationships in Incrementally Trained Interpretable Models
- 设计多模态概念来建模概念-类别间复杂关系,不增加参数量。
- 在部分任务上分类性能超现有方法2倍以上,且能有效防止概念遗忘。
- 适合需要持续学习且注重可解释性的实际应用场景。
概念驱动的方法在监督学习中为可解释神经网络提供了有前景的方向。然而,大多数增量学习研究要么假设概念集恒定,要么假设每个学习阶段使用独立概念集。本文研究更现实的动态场景:新类别可能复用旧概念并引入新概念。我们发现概念与类别形成复杂关系网,易受退化影响,需在跨阶段中同时保留与扩展。现有方法虽采用防灾难性遗忘策略,仍无法同步维护概念、类别及概念-类别关系。为此,我们提出新方法 MuCIL,利用多模态概念进行分类,不增加训练参数。这些概念与自然语言描述对齐,具备天生可解释性。大量实验表明,该方法在分类性能上达到当前最优,某些情况下超越基准超过2倍。此外,模型支持概念干预,能定位输入图像中的视觉概念,提供事后解释。
原文摘要 · Abstract (English)
Concept-based methods have emerged as a promising direction to develop interpretable neural networks in standard supervised settings. However, most works that study them in incremental settings assume either a static concept set across all experiences or assume that each experience relies on a distinct set of concepts. In this work, we study concept-based models in a more realistic, dynamic setting where new classes may rely on older concepts in addition to introducing new concepts themselves. We show that concepts and classes form a complex web of relationships, which is susceptible to degradation and needs to be preserved and augmented across experiences. We introduce new metrics to show that existing concept-based models cannot preserve these relationships even when trained using methods to prevent catastrophic forgetting, since they cannot handle forgetting at concept, class, and concept-class relationship levels simultaneously. To address these issues, we propose a novel method - MuCIL - that uses multimodal concepts to perform classification without increasing the number of trainable parameters across experiences. The multimodal concepts are aligned to concepts provided in natural language, making them interpretable by design. Through extensive experimentation, we show that our approach obtains state-of-the-art classification performance compared to other concept-based models, achieving over 2$\times$ the classification performance in some cases. We also study the ability of our model to perform interventions on concepts, and show that it can localize visual concepts in input images, providing post-hoc interpretations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。