提出可跨模态的可信概念瓶颈模型,解决解释性与准确性的矛盾。
Towards Faithful Multimodal Concept Bottleneck Models
- 用可微分泄漏损失与柯尔莫哥洛夫网络头联合优化概念检测与信息泄露。
- 在多模态数据上实现准确率、概念检测和泄漏抑制的最佳平衡。
- 适配图像、文本及纯文本数据,跨模态通用性强。
概念瓶颈模型(CBMs)通过人类可理解的概念层进行预测,具有良好的可解释性。尽管在视觉和自然语言处理领域已有研究,但多模态场景下的探索仍不足。为确保解释的可信性,CBMs需满足两个条件:概念需被正确检测,且概念表示仅包含目标语义,不携带任务相关或概念间干扰信息,即避免泄漏。现有方法将概念检测与泄漏缓解视为独立问题,常以牺牲预测准确率为代价提升解释性。本文提出f-CBM,基于视觉-语言骨干网络,通过两种互补策略联合优化:可微分泄漏损失以减少泄漏,柯尔莫哥洛夫-阿诺德网络预测头提升概念检测能力。实验表明,f-CBM在任务准确率、概念检测与泄漏抑制之间取得最佳权衡,并能无缝应用于图像、文本或纯文本数据集,具备跨模态通用性。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) are interpretable models that route predictions through a layer of human-interpretable concepts. While widely studied in vision and, more recently, in NLP, CBMs remain largely unexplored in multimodal settings. For their explanations to be faithful, CBMs must satisfy two conditions: concepts must be properly detected, and concept representations must encode only their intended semantics, without smuggling extraneous task-relevant or inter-concept information into final predictions, a phenomenon known as leakage. Existing approaches treat concept detection and leakage mitigation as separate problems, and typically improve one at the expense of predictive accuracy. In this work, we introduce f-CBM, a faithful multimodal CBM framework built on a vision-language backbone that jointly targets both aspects through two complementary strategies: a differentiable leakage loss to mitigate leakage, and a Kolmogorov-Arnold Network prediction head that provides sufficient expressiveness to improve concept detection. Experiments demonstrate that f-CBM achieves the best trade-off between task accuracy, concept detection, and leakage reduction, while applying seamlessly to both image and text or text-only datasets, making it versatile across modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。