让图像分类决策可解释,定位到具体区域和概念
Spatially Grounded Concept-Based Image Classification
- 将图像切分为概念引导的区域,用注意力聚合局部证据
- 在Waterbirds数据集上最差组准确率提升至72.0%,CIFAR-100达85.3%
- 无需额外模块即可同时生成预测与解释,适合需要透明决策的场景
深度神经网络虽能取得高准确率,但依赖的判别依据难以检验或与任务目标不一致。概念瓶颈模型(CBMs)虽能揭示人类可理解的概念,但多数将其视为全局属性,未说明局部证据如何汇聚成决策。我们提出 extbf{SEG-MIL-CBM},一种空间对齐的CBM,将每张图像分解为概念引导的区域,并通过注意力机制聚合片段级概念证据进行分类。相同的片段证据同时用于预测与解释,无需额外后处理归因模块即可揭示支持预测得分的具体区域与概念。在对比的CBM基线中,SEG-MIL-CBM将Waterbirds最差组准确率从65.1%提升至72.0%,在Pawrious上达到87.4%最差组准确率,标准识别任务表现仍具竞争力,在CIFAR-100上实现85.3%最佳CBM准确率。在CUB上的片段忠实度实验表明,其学习的片段排序表现匹配或优于对比控制方法。
原文摘要 · Abstract (English)
Deep neural networks can achieve high accuracy while relying on evidence that is hard to inspect or misaligned with the intended task. Concept Bottleneck Models (CBMs) expose human-interpretable concepts, but most treat concepts as global attributes and do not show how localized evidence is aggregated into a decision. We propose \textbf{SEG-MIL-CBM}, a spatially grounded CBM that decomposes each image into concept-guided regions and classifies it by attention-based aggregation of segment-level concept evidence. The same segment evidence terms form the prediction and the explanation, exposing which regions and concepts support the predicted logit without a separate post-hoc attribution module. Among evaluated CBM-family baselines, SEG-MIL-CBM improves Waterbirds worst-group accuracy from $65.1\%$ to $72.0\%$, reaches $87.4\%$ worst-group accuracy on Pawrious, remains competitive on standard recognition, and attains the best CBM accuracy on CIFAR-100 ($85.3\%$). Segment-level faithfulness experiments on CUB further show that its learned segment ranking matches or improves over evaluated segment-ranking controls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。