用视觉原型让概念模型可验证,提升可解释性与可控性
Prototype-Grounded Concept Models for Verifiable Concept Alignment

- 用图像局部特征作为概念的视觉原型,实现概念可追溯
- 预测性能与顶尖概念模型相当,但可解释性显著提升
- 支持人工干预原型来修正概念误对齐,适合需要透明决策的场景
概念瓶颈模型(CBMs)通过人类可理解的概念结构化深度学习预测,但无法验证所学概念是否符合人类预期,损害了可解释性。本文提出原型锚定概念模型(PGCMs),将概念基于学习到的视觉原型——即能明确支撑概念的图像局部区域。这种锚定使概念语义可直接检验,并支持在原型层面进行针对性的人工干预以纠正偏差。实证表明,PGCMs在预测性能上与当前最先进的CBMs相当,同时大幅提升了透明度、可解释性和可干预性。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) aim to improve interpretability in Deep Learning by structuring predictions through human-understandable concepts, but they provide no way to verify whether learned concepts align with the human's intended meaning, hurting interpretability. We introduce Prototype-Grounded Concept Models (PGCMs), which ground concepts in learned visual prototypes: image parts that serve as explicit evidence for the concepts. This grounding enables direct inspection of concept semantics and supports targeted human intervention at the prototype level to correct misalignments. Empirically, PGCMs achieve similar predictive performance as state-of-the-art CBMs while substantially improving transparency, interpretability, and intervenability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。