arXiv:2602.21944cs.CV2026-02

通过图融合提升多视角眼底图像的糖尿病视网膜病变分级准确率

Learning to Fuse and Reconstruct Multi-View Graphs for Diabetic Retinopathy Grading

  • 构建多视图图结构,分离共享与视角特异性特征
  • 在MFIDDR数据集上达到92.1%准确率,优于现有方法
  • 适合医学图像分析与多模态学习研究者参考

糖尿病视网膜病变(DR)是全球致盲的主要原因之一,早期精准分级对及时干预至关重要。当前临床实践采用多视角眼底图像以扩大视野覆盖,推动深度学习方法探索多视角学习在DR分级中的潜力。然而,现有方法在融合多视角图像时常忽视视间相关性,未能充分挖掘来自同一患者的不同视角间的内在一致性。本文提出端到端的多视角图融合框架MVGFDR,其核心为新型多视角图融合(MVGF)模块,可显式解耦共享与视角特异性视觉特征。具体包括:(1)多视角图初始化,通过残差引导连接构建视觉图,并以离散余弦变换(DCT)系数作为频域锚点;(2)多视角图融合,基于频域相关性选择性整合多视角图节点,捕获互补的视角特异性信息;(3)掩码跨视角重建,利用跨视角共享信息的掩码重建,促进视角不变表示学习。在目前最大的多视角眼底图像数据集MFIDDR上的大量实验表明,所提方法在糖尿病视网膜病变分级任务中显著优于现有最先进方法。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is one of the leading causes of vision loss worldwide, making early and accurate DR grading critical for timely intervention. Recent clinical practices leverage multi-view fundus images for DR detection with a wide coverage of the field of view (FOV), motivating deep learning methods to explore the potential of multi-view learning for DR grading. However, existing methods often overlook the inter-view correlations when fusing multi-view fundus images, failing to fully exploit the inherent consistency across views originating from the same patient. In this work, we present MVGFDR, an end-to-end Multi-View Graph Fusion framework for DR grading. Different from existing methods that directly fuse visual features from multiple views, MVGFDR is equipped with a novel Multi-View Graph Fusion (MVGF) module to explicitly disentangle the shared and view-specific visual features. Specifically, MVGF comprises three key components: (1) Multi-view Graph Initialization, which constructs visual graphs via residual-guided connections and employs Discrete Cosine Transform (DCT) coefficients as frequency-domain anchors; (2) Multi-view Graph Fusion, which integrates selective nodes across multi-view graphs based on frequency-domain relevance to capture complementary view-specific information; and (3) Masked Cross-view Reconstruction, which leverages masked reconstruction of shared information across views to facilitate view-invariant representation learning. Extensive experimental results on MFIDDR, by far the largest multi-view fundus image dataset, demonstrate the superiority of our proposed approach over existing state-of-the-art approaches in diabetic retinopathy grading.

医学图像图神经网络多视角学习糖尿病视网膜病变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。