研究概念瓶颈模型中噪声标注的影响,提出稳定训练与推理修正方法。
An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
- 用尖锐感知优化稳定易受噪声影响的概念学习
- 仅修正高不确定性概念,减少90%以上性能损失
- 适合关注可解释性与鲁棒性平衡的研究者
概念瓶颈模型(CBMs)通过将预测分解为人类可理解的概念来保证可解释性。然而,用于训练的标注常含噪声,其影响尚未被充分理解。本研究首次系统分析了CBMs中的噪声问题,发现适度噪声会同时损害预测性能、可解释性和干预有效性。分析识别出一小部分概念对噪声极度敏感,其准确率下降远超平均差距,且是性能损失的主要来源。为此,我们提出两阶段框架:训练阶段采用尖锐感知最小化(SAM)以稳定噪声敏感概念的学习;推理阶段基于预测熵排序概念,仅修正最不确定项,利用不确定性作为敏感性的代理。理论分析与大量消融实验揭示了SAM提升鲁棒性的机制及不确定性识别敏感概念的可靠性,为在噪声环境下保持可解释性与鲁棒性提供了原则性依据。
原文摘要 · Abstract (English)
Concept bottleneck models (CBMs) ensure interpretability by decomposing predictions into human interpretable concepts. Yet the annotations used for training CBMs that enable this transparency are often noisy, and the impact of such corruption is not well understood. In this study, we present the first systematic study of noise in CBMs and show that even moderate corruption simultaneously impairs prediction performance, interpretability, and the intervention effectiveness. Our analysis identifies a susceptible subset of concepts whose accuracy declines far more than the average gap between noisy and clean supervision and whose corruption accounts for most performance loss. To mitigate this vulnerability we propose a two-stage framework. During training, sharpness-aware minimization stabilizes the learning of noise-sensitive concepts. During inference, where clean labels are unavailable, we rank concepts by predictive entropy and correct only the most uncertain ones, using uncertainty as a proxy for susceptibility. Theoretical analysis and extensive ablations elucidate why sharpness-aware training confers robustness and why uncertainty reliably identifies susceptible concepts, providing a principled basis that preserves both interpretability and resilience in the presence of noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。