arXiv:2512.13008cs.CV2025-12被引 1

用文本引导弱监督方法实现糖尿病视网膜病变的可解释定位与分级。

TWLR: Text-Guided Weakly-Supervised Lesion Localization and Severity Regression for Explainable Diabetic Retinopathy Grading

  • 通过视觉语言模型融合眼科知识,联合完成分级与病灶分类。
  • 迭代回归框架实现无像素标注的病灶定位,准确率超基线。
  • 可视化疾病向健康状态演变过程,适合临床可解释性需求。

精准的医学图像分析能显著辅助临床诊断,但其效果依赖高质量专家标注。获取医学图像尤其是眼底图像的像素级标签仍成本高昂且耗时。尽管深度学习在医学影像领域取得成功,但缺乏可解释性限制了其临床应用。为此,我们提出TWLR,一种两阶段可解释性糖尿病视网膜病变(DR)评估框架。第一阶段,利用视觉语言模型将领域专有眼科知识融入文本嵌入,联合完成DR分级与病灶分类,有效关联语义医学概念与视觉特征。第二阶段引入基于弱监督语义分割的迭代严重程度回归框架。通过迭代优化生成的病灶显著性图驱动渐进式修复机制,系统性消除病理特征,使疾病严重程度逐步向健康眼底形态退化。该方法兼具双重优势:无需像素级监督实现精确病灶定位,并提供疾病向健康转化的可解释可视化。在FGADR、DDR及私有数据集上的实验表明,TWLR在DR分类与病灶分割任务中均达到具有竞争力的性能,为自动化视网膜图像分析提供了更可解释、注释效率更高的解决方案。

原文摘要 · Abstract (English)

Accurate medical image analysis can greatly assist clinical diagnosis, but its effectiveness relies on high-quality expert annotations Obtaining pixel-level labels for medical images, particularly fundus images, remains costly and time-consuming. Meanwhile, despite the success of deep learning in medical imaging, the lack of interpretability limits its clinical adoption. To address these challenges, we propose TWLR, a two-stage framework for interpretable diabetic retinopathy (DR) assessment. In the first stage, a vision-language model integrates domain-specific ophthalmological knowledge into text embeddings to jointly perform DR grading and lesion classification, effectively linking semantic medical concepts with visual features. The second stage introduces an iterative severity regression framework based on weakly-supervised semantic segmentation. Lesion saliency maps generated through iterative refinement direct a progressive inpainting mechanism that systematically eliminates pathological features, effectively downgrading disease severity toward healthier fundus appearances. Critically, this severity regression approach achieves dual benefits: accurate lesion localization without pixel-level supervision and providing an interpretable visualization of disease-to-healthy transformations. Experimental results on the FGADR, DDR, and a private dataset demonstrate that TWLR achieves competitive performance in both DR classification and lesion segmentation, offering a more explainable and annotation-efficient solution for automated retinal image analysis.

可解释性弱监督眼底病变视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。