arXiv:2508.15251eess.IVcs.AI2025-08

用教师模型指导小模型,让医疗影像分类更高效且可解释。

Explainable Knowledge Distillation for Efficient Medical Image Classification

  • 用大模型指导小模型,融合真实标签与软标签提升性能。
  • 学生模型参数减少,推理速度更快,准确率仍保持高位。
  • 可视化分析模型关注区域,适合临床可信AI应用。

本研究系统探索了用于新冠和肺癌分类的医学影像知识蒸馏框架,基于胸部X光(CXR)图像。采用高容量教师模型(VGG19及轻量级Vision Transformers,如Visformer-S、AutoFormer-V2-T),指导由OFA-595超网络衍生的紧凑型硬件感知学生模型训练。方法结合真实标签与教师模型的软目标,实现精度与计算效率的平衡。在两个基准数据集COVID-QU-Ex和LCS25000上验证,涵盖新冠、健康、非新冠肺炎、肺癌及结肠癌等多类。通过Score-CAM可视化分析模型空间注意力分布,揭示师生网络的决策依据。结果表明,蒸馏后学生模型显著降低参数量与推理时间,同时保持高分类性能,适用于资源受限的临床场景。本工作强调了效率与可解释性结合对构建可信医疗AI的关键价值。

原文摘要 · Abstract (English)

This study comprehensively explores knowledge distillation frameworks for COVID-19 and lung cancer classification using chest X-ray (CXR) images. We employ high-capacity teacher models, including VGG19 and lightweight Vision Transformers (Visformer-S and AutoFormer-V2-T), to guide the training of a compact, hardware-aware student model derived from the OFA-595 supernet. Our approach leverages hybrid supervision, combining ground-truth labels with teacher models' soft targets to balance accuracy and computational efficiency. We validate our models on two benchmark datasets: COVID-QU-Ex and LCS25000, covering multiple classes, including COVID-19, healthy, non-COVID pneumonia, lung, and colon cancer. To interpret the spatial focus of the models, we employ Score-CAM-based visualizations, which provide insight into the reasoning process of both teacher and student networks. The results demonstrate that the distilled student model maintains high classification performance with significantly reduced parameters and inference time, making it an optimal choice in resource-constrained clinical environments. Our work underscores the importance of combining model efficiency with explainability for practical, trustworthy medical AI solutions.

知识蒸馏医疗影像可解释性模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。