让概念瓶颈模型可编辑,无需重训就能改数据、调概念。
Controllable Concept Bottleneck Models
- 通过影响函数推导出闭式近似,实现无需重训的编辑。
- 支持概念标签、概念、数据三级精细编辑,覆盖增删改。
- 适用于需持续维护的可信AI系统,如医疗、金融场景。
概念瓶颈模型(CBMs)因其可通过人类可理解的概念层解释预测过程而受到广泛关注。然而,以往研究多集中于静态场景,假设数据和概念固定且无噪声。在真实应用中,模型需持续维护:常需移除错误或敏感数据(遗忘)、修正误标注概念,或加入新样本(增量学习)以适应动态环境。因此,在不重新训练的前提下实现高效可编辑的CBM仍是重大挑战,尤其在大规模应用中。为此,我们提出可控概念瓶颈模型(CCBMs)。具体而言,CCBMs支持概念-标签级、概念级和数据级三种粒度的模型编辑,后者涵盖数据删除与添加。CCBMs基于影响函数推导出数学上严谨的闭式近似,避免了重新训练。实验表明,所提方法在效率与适应性上表现优异,证实其在构建动态、可信的CBM中的实际价值。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) have garnered much attention for their ability to elucidate the prediction process through a human-understandable concept layer. However, most previous studies focused on static scenarios where the data and concepts are assumed to be fixed and clean. In real-world applications, deployed models require continuous maintenance: we often need to remove erroneous or sensitive data (unlearning), correct mislabeled concepts, or incorporate newly acquired samples (incremental learning) to adapt to evolving environments. Thus, deriving efficient editable CBMs without retraining from scratch remains a significant challenge, particularly in large-scale applications. To address these challenges, we propose Controllable Concept Bottleneck Models (CCBMs). Specifically, CCBMs support three granularities of model editing: concept-label-level, concept-level, and data-level, the latter of which encompasses both data removal and data addition. CCBMs enjoy mathematically rigorous closed-form approximations derived from influence functions that obviate the need for retraining. Experimental results demonstrate the efficiency and adaptability of our CCBMs, affirming their practical value in enabling dynamic and trustworthy CBMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。