arXiv:2605.09242eess.IVcs.CV2026-05

用视觉语言模型提升糖尿病视网膜病变分级准确率

Cross-Modal Semantic-Enhanced Diffusion Framework for Diabetic Retinopathy Grading

论文配图:Cross-Modal Semantic-Enhanced Diffusion Framework for Diabetic Retinopathy Grading
图 1 · 摘自论文原文
  • 引入领域专用视觉语言模型作为语义引导,通过低秩适配微调
  • 跨模态融合图像与分级文本特征,生成联合表征用于扩散模型
  • 在APTOS 2019数据集上达到87.5%准确率,优于现有方法

糖尿病视网膜病变(DR)自动分级面临多重挑战:细粒度病变模式在不同等级间差异细微,异构成像设备和采集条件导致的数据分布偏差,以及纯视觉方法难以利用临床语义知识。本文提出CLIP引导的语义扩散框架(CGSD),将视觉-语言预训练与扩散概率建模相结合。采用专为DR分级定制的视觉-语言模型作为语义引导模块,并通过低秩适配(LoRA)进行领域自适应,仅用少量可训练参数有效弥合预训练模型与目标数据集间的分布差距。在此基础上,通过计算图像特征与各等级文本描述特征的点积,构建跨模态语义条件向量,作为扩散去噪网络的条件信号,替代现有方法中结构复杂的双分支视觉先验。在APTOS 2019数据集上的实验表明,该方法实现87.5%的准确率和0.731的宏平均F1分数,优于多种代表性方法。消融实验证明了各模块的独立贡献。

原文摘要 · Abstract (English)

Automated grading of diabetic retinopathy (DR) faces several critical challenges: subtle inter-grade visual distinctions in fine-grained lesion patterns, distributional discrepancies induced by heterogeneous imaging devices and acquisition conditions, and the inherent inability of purely visual approaches to exploit clinical semantic knowledge. In this paper, we propose CLIP-Guided Semantic Diffusion (CGSD), a DR grading framework that synergistically integrates vision-language pretraining with diffusion probabilistic modeling. We adopt a domain-specific vision-language model tailored for DR grading as the semantic guidance module and adapt it to the target domain via Low-Rank Adaptation (LoRA), effectively bridging the distributional gap between the pretrained model and the target dataset with only a minimal number of trainable parameters. Building on this foundation, we construct a cross-modal semantic conditioning vector by computing the dot product between image features and the text description features of each DR grade, yielding a joint representation that simultaneously encodes visual content and clinical-grade semantics. This vector serves as the conditioning signal for the diffusion denoising network, replacing the structurally complex dual-branch visual prior employed in existing diffusion-based classification methods. Experiments on the APTOS 2019 dataset demonstrate that the proposed approach achieves an accuracy of 87.5% and a macro-averaged F1 score of 0.731, outperforming a variety of representative methods. Ablation studies further validate the independent contribution of each constituent module.

糖尿病视网膜病变扩散模型多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。