不同模型架构在微调后解释逻辑变化,影响医疗影像诊断可信度。
When Fine-Tuning Changes the Evidence: Architecture-Dependent Semantic Drift in Chest X-Ray Explanations
- 对比三种模型在微调前后解释图的结构变化,发现证据依赖关系重组。
- 准确率稳定但解释区域重分布,尤其在解剖重叠区域差异显著。
- 不同归因方法结果相反,说明解释稳定性取决于模型与方法交互。
迁移学习结合微调在医学图像分类中广泛应用,且能持续提升诊断性能。但在具有视觉特征重叠的多类别任务中,准确率提升并不保证预测所依赖的视觉证据保持稳定。本文将语义漂移定义为从迁移学习到全微调过程中,支撑模型预测的归因结构系统性变化,反映潜在视觉推理机制的转移,尽管分类性能保持稳定。基于五类胸部X光片任务,评估DenseNet201、ResNet50V2和InceptionV3在两阶段训练下的表现,使用无参考指标量化归因图的空间定位与结构一致性。结果显示:粗略解剖定位保持稳定,而交叠区域的交并比(IoU)揭示了显著的架构依赖型证据结构重组。此外,在预测性能收敛后,不同归因方法(LayerCAM与GradCAM++)的稳定性排序出现反转,表明解释稳定性是架构、优化阶段与归因目标之间复杂互动的结果。
原文摘要 · Abstract (English)
Transfer learning followed by fine-tuning is widely adopted in medical image classification due to consistent gains in diagnostic performance. However, in multi-class settings with overlapping visual features, improvements in accuracy do not guarantee stability of the visual evidence used to support predictions. We define semantic drift as systematic changes in the attribution structure supporting a model's predictions between transfer learning and full fine-tuning, reflecting potential shifts in underlying visual reasoning despite stable classification performance. Using a five-class chest X-ray task, we evaluate DenseNet201, ResNet50V2, and InceptionV3 under a two-stage training protocol and quantify drift with reference-free metrics capturing spatial localization and structural consistency of attribution maps. Across architectures, coarse anatomical localization remains stable, while overlap IoU reveals pronounced architecture-dependent reorganization of evidential structure. Beyond single-method analysis, stability rankings can reverse across LayerCAM and GradCAM++ under converged predictive performance, establishing explanation stability as an interaction between architecture, optimization phase, and attribution objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。