用视觉语言模型让糖尿病视网膜病变分级结果可解释,提升临床可信度。
From Pixels to Explanations: Interpretable Diabetic Retinopathy Grading with CNN-Transformer Ensembles, Visual Explainability and Vision-Language Models
- 融合CNN与Transformer,通过加权软投票提升分级一致性。
- 加权软投票使跨折平均QWK达0.934,优于其他集成策略。
- 结合梯度热图与视觉语言模型生成可读的诊断理由,适合医生参考。
糖尿病视网膜病变(DR)筛查的准确性依赖于对病情严重程度的正确分级;然而,许多深度学习分类器在临床上难以解释。本研究提出一种结合强判别模型与多模态解释的方法,将视网膜图像像素转化为临床可理解的输出。在APTOS 2019基准上,采用分层五折交叉验证,评估六种代表性基于CNN和Transformer的骨干网络。比较了硬投票、加权软投票、堆叠等集成策略,并探索了混合级融合以利用各等级优势。对于可解释性,使用Grad-CAM++生成视觉归因图,并在保守提示约束下,通过视觉语言模型(VLM)生成短文本理由。现代CNN骨干(ResNet-50和ConvNeXt-Tiny)表现最佳,交叉验证下加权κ值分别达0.919和0.914。加权软投票在各折中最为稳定(0.934 ± 0.017)。混合级融合表现相当但未在配对折叠比较中显著优于标准融合(霍尔姆校正p ≥ 1.000)。VLM生成的理由整体与等级一致,定量上存在临床完整性与模板语义相似性之间的权衡(覆盖率0.700,BERTScore 0.072),图像-文本对齐能力良好(CLIPScore约0.34)。
原文摘要 · Abstract (English)
The quality of diabetic retinopathy (DR) screening relies on the ability to correctly grade severity; however, many deep-learning (DL) classifiers cannot be easily interpreted in the clinical context. This study presents a methodology that combines strong discriminative models with multimodal explanations, converting retinal pixels into clinically interpretable outputs. Using the APTOS 2019 benchmark, we evaluated six representative CNN- and transformer-based backbones under a controlled protocol with stratified five-fold cross-validation. We then compared ensembling strategies (hard voting, weighted soft voting, stacking) and investigated a hybrid class-level fusion variant to exploit grade-specific advantages. For interpretability, we produced Grad-CAM++ visual attribution maps and short textual rationales using vision-language models (VLMs) conditioned on the fundus image and classifier outputs under conservative prompting constraints. Modern CNN backbones (ResNet-50 and ConvNeXt-Tiny) provided the strongest single-model baselines, with cross-validated QWK up to 0.919 and 0.914, respectively. Ensembling improved ordinal agreement, and weighted soft voting was the most consistent across folds (QWK 0.934 +/- 0.017). Hybrid class-level fusion was competitive but did not yield a statistically reliable improvement over standard fusion in paired fold comparisons (Holm-adjusted p >= 1.000). For explanation quality, Grad-CAM++ offered plausible but coarse localization, and VLM rationales were generally grade-consistent. Quantitatively, VLM variants showed a trade-off between clinical completeness and template-level semantic similarity (coverage 0.700 vs. BERTScore 0.072), while image-text alignment was comparable (CLIPScore approximately 0.34).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。