arXiv:2510.26315cs.CV2025-10被引 1

融合CNN与ViT优势,用证据理论提升糖尿病视网膜病变分级准确率。

A Hybrid Framework Bridging CNN and ViT based on Theory of Evidence for Diabetic Retinopathy Grading

  • 基于证据理论将CNN与ViT特征转为支持证据,动态融合。
  • 在两个公开数据集上达到更优分类精度,优于现有方法。
  • 融合过程可解释,适合医疗诊断场景需求。

糖尿病视网膜病变(DR)是中老年人视力丧失的主要原因,严重影响生活质量与心理健康。为提高临床筛查效率并实现早期诊断,已有多种基于卷积神经网络(CNN)或视觉变压器(ViT)的自动化诊断系统。然而,由于单一结构的固有局限性,现有方法性能已逼近瓶颈。进一步提升的关键在于融合不同骨干网络的优势——即利用CNN的局部特征提取能力与ViT的全局建模能力。为此,本文提出一种基于证据理论的新型混合框架,通过深度证据网络将不同骨干网络的特征转化为支持证据,并据此生成聚合意见,自适应调节融合模式,从而提升模型性能。我们在两个公开的DR分级数据集上评估了该方法,实验结果表明,所提混合模型不仅在分类准确率上超越当前最优框架,还提供了特征融合与决策过程的优秀可解释性。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is a leading cause of vision loss among middle-aged and elderly people, which significantly impacts their daily lives and mental health. To improve the efficiency of clinical screening and enable the early detection of DR, a variety of automated DR diagnosis systems have been recently established based on convolutional neural network (CNN) or vision Transformer (ViT). However, due to the own shortages of CNN / ViT, the performance of existing methods using single-type backbone has reached a bottleneck. One potential way for the further improvements is integrating different kinds of backbones, which can fully leverage the respective strengths of them (\emph{i.e.,} the local feature extraction capability of CNN and the global feature capturing ability of ViT). To this end, we propose a novel paradigm to effectively fuse the features extracted by different backbones based on the theory of evidence. Specifically, the proposed evidential fusion paradigm transforms the features from different backbones into supporting evidences via a set of deep evidential networks. With the supporting evidences, the aggregated opinion can be accordingly formed, which can be used to adaptively tune the fusion pattern between different backbones and accordingly boost the performance of our hybrid model. We evaluated our method on two publicly available DR grading datasets. The experimental results demonstrate that our hybrid model not only improves the accuracy of DR grading, compared to the state-of-the-art frameworks, but also provides the excellent interpretability for feature fusion and decision-making.

医学图像特征融合证据理论糖尿病视网膜病变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。