用自监督与人工引导结合提升Transformer在胸片诊断中的可解释性与泛化能力
Hybrid Explanation-Guided Learning for Transformer-Based Chest X-Ray Diagnosis
- 融合自监督与人工引导的注意力对齐机制,减少模型偏差
- 在胸片分类任务中准确率超越现有最优方法,且注意力图更贴近医生判断
- 适用于需高可信度医疗AI的场景,尤其关注模型可解释性
基于Transformer的深度学习模型在医学影像中凭借注意力机制展现出优异性能,但易学习虚假关联,导致偏差并限制泛化能力。尽管人类与AI注意力对齐可缓解此问题,却常依赖昂贵的人工标注。本文提出混合解释引导学习(H-EGL)框架,结合自监督与人工引导约束以增强注意力对齐并提升泛化能力。其自监督部分利用类别特异性注意力,无需强先验假设,具备更强鲁棒性与灵活性。在使用视觉Transformer(ViT)进行胸片分类的任务中,H-EGL优于两种前沿的解释引导学习(EGL)方法,表现更优的分类准确率与泛化能力,并生成更符合人类专家判断的注意力热力图。
原文摘要 · Abstract (English)
Transformer-based deep learning models have demonstrated exceptional performance in medical imaging by leveraging attention mechanisms for feature representation and interpretability. However, these models are prone to learning spurious correlations, leading to biases and limited generalization. While human-AI attention alignment can mitigate these issues, it often depends on costly manual supervision. In this work, we propose a Hybrid Explanation-Guided Learning (H-EGL) framework that combines self-supervised and human-guided constraints to enhance attention alignment and improve generalization. The self-supervised component of H-EGL leverages class-distinctive attention without relying on restrictive priors, promoting robustness and flexibility. We validate our approach on chest X-ray classification using the Vision Transformer (ViT), where H-EGL outperforms two state-of-the-art Explanation-Guided Learning (EGL) methods, demonstrating superior classification accuracy and generalization capability. Additionally, it produces attention maps that are better aligned with human expertise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。