arXiv:2604.19323cs.LGcs.CV2026-04

发现皮肤镜数据中16.4%概念组合存在诊断矛盾,导致模型准确率被硬性限制。

Concept Inconsistency in Dermoscopic Concept Bottleneck Models: A Rough-Set Analysis of the Derm7pt Dataset

论文配图:Concept Inconsistency in Dermoscopic Concept Bottleneck Models: A Rough-Set Analysis of the Derm7pt Dataset
图 1 · 摘自论文原文
  • 用粗糙集理论分析皮肤镜数据,识别概念层不一致问题
  • 305种概念组合中有50种不一致,理论准确率上限为92.1%
  • 提出过滤策略构建无矛盾数据集Derm7pt+,支持可靠模型评估

概念瓶颈模型(CBM)通过临床可解释的概念层进行预测,其性能依赖于概念与标签的一致性。当数据集中存在概念层面的不一致时,相同概念配置对应不同诊断结果,形成无法解决的瓶颈,导致模型准确率存在理论上限。本文对Derm7pt皮肤镜基准数据集应用粗糙集理论,系统刻画了该不一致性的范围与临床结构。在由7项皮肤镜标准构成的305个唯一概念组合中,有50个(16.4%)存在不一致,涉及306张图像(占数据集30.3%),由此推导出基于硬概念的CBM理论准确率上限为92.1%,与主干网络或训练策略无关。此外,我们分析了冲突严重度分布,识别出造成边界模糊的关键临床特征,并评估两种过滤策略对数据集构成与可解释性的影响。对称移除所有边界区域图像后得到新数据集Derm7pt+,包含705张完全一致图像,具备理想分类质量且无硬性准确率上限。基于此数据集,我们在19种主干网络(EfficientNet、DenseNet、ResNet、Wide ResNet系列)上测试硬性CBM,对称过滤下EfficientNet-B5表现最佳(标签F1 0.85,准确率0.90,概念准确率0.70);不对称过滤下EfficientNet-B7在四项指标上领先,标签F1达0.82,概念准确率0.70。这些结果为皮肤镜领域概念一致性的CBM评估提供了可复现基准。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) route predictions exclusively through a clinically grounded concept layer, binding interpretability to concept-label consistency. When a dataset contains concept-level inconsistencies, identical concept profiles mapped to conflicting diagnosis labels create an unresolvable bottleneck that imposes a hard ceiling on achievable accuracy. In this paper, we apply rough set theory to the Derm7pt dermoscopy benchmark and characterize the full extent and clinical structure of this inconsistency. Among 305 unique concept profiles formed by the 7 dermoscopic criteria of the 7-point melanoma checklist, 50 (16.4%) are inconsistent, spanning 306 images (30.3% of the dataset). This yields a theoretical accuracy ceiling of 92.1%, independent of backbone architecture or training strategy for CBMs that exclusively operate with hard concepts. In addition, we characterize the conflict-severity distribution, identify the clinical features most responsible for boundary ambiguity, and evaluate two filtering strategies with quantified effects on dataset composition and CBM interpretability. Symmetric removal of all boundary-region images yields Derm7pt+, a fully consistent benchmark subset of 705 images with perfect quality of classification and no hard accuracy ceiling. Building on this filtered dataset, we present a hard CBM evaluated across 19 backbone architectures from the EfficientNet, DenseNet, ResNet, and Wide ResNet families. Under symmetric filtering, explored for completeness, EfficientNet-B5 achieves the best label F1 score (0.85) and label accuracy (0.90) on the held-out test set, with a concept accuracy of 0.70. Under asymmetric filtering, EfficientNet-B7 leads across all four metrics, reaching a label F1 score of 0.82 and concept accuracy of 0.70. These results establish reproducible baselines for concept-consistent CBM evaluation on dermoscopic data.

概念瓶颈皮肤镜粗糙集可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。