arXiv:2606.15129cs.CVcs.AI2026-06

用OCT信息增强眼底彩照模型,让普通筛查更准

EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining

论文配图:EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining
图 1 · 摘自论文原文
  • 通过成对眼底彩照与OCT图像预训练,学习带深度信息的彩照特征
  • 在15个任务中表现优于或持平主流模型,尤其在黄斑和视神经病变上提升显著
  • 仅需眼底彩照推理,适合临床大规模筛查应用

彩色眼底摄影(CFP)是大规模视网膜筛查的主要手段,但受限于缺乏深度结构信息;光学相干断层扫描(OCT)可提供该信息,却难以在人群规模上普及。我们提出EyeMVP,一种基于成对CFP--OCT预训练的跨模态视网膜基础模型,能在仅使用CFP进行推理的前提下,学习到受OCT启发的CFP表示。模型在来自8家医院、涵盖112,642名患者、共674,893组同眼同日的CFP--OCT数据上预训练,采用跨模态掩码重建机制,将OCT相关监督融入CFP特征,并结合源约束跨注意力与基于CFP的结构掩码,以适应横断面OCT与俯视图CFP之间的几何非对齐问题。在15个涵盖分类与分割的任务设置下,无论全数据还是少样本场景,EyeMVP均达到或超过现有代表性视网膜基础模型的表现,尤其在黄斑和视神经相关任务上持续提升;其在黄斑水肿(AUROC 0.923)和病理性近视黄斑裂孔(AUROC 0.867)两项在传统CFP中难以分辨的疾病上表现突出。探索性阅片研究显示,EyeMVP在黄斑水肿上超越初级与中级眼科医生,但不及资深医生;而在病理性近视黄斑裂孔上则全面超越所有医生群体。结果表明,跨模态重建能有效将OCT相关信息注入CFP表示,为提升基于CFP的筛查能力提供了可行路径。

原文摘要 · Abstract (English)

Color fundus photography (CFP) is the mainstay of large-scale retinal screening, but its diagnostic capacity is limited by the lack of depth-resolved structure, which optical coherence tomography (OCT) provides yet is less accessible at population scale. We present EyeMVP, a cross-modal retinal foundation model that uses paired CFP--OCT pretraining to learn OCT-informed CFP representations while requiring only CFP at inference. Pretrained on 674,893 same-eye same-day CFP--OCT triples from 112,642 patients across eight hospitals, EyeMVP uses cross-modal masked reconstruction to enrich CFP features with OCT-associated supervision, and combines source-constrained cross-attention with CFP-derived structural masks to accommodate the non-aligned geometry of en-face CFP and cross-sectional OCT. Across 15 dataset-level settings spanning classification and segmentation, under both full-data and few-shot regimes, EyeMVP performs on par with or better than representative retinal foundation models, with consistent gains on macular and optic-nerve tasks; it attains AUROCs of 0.923 for macular edema and 0.867 for myopic macular schisis, two conditions poorly resolved in CFP. In an exploratory reader study, EyeMVP surpasses junior and intermediate ophthalmologists but not seniors on macular edema, while exceeding all groups on myopic macular schisis. These results indicate that cross-modal reconstruction can enrich CFP representations with OCT-associated supervision, offering a practical route to stronger CFP-based screening.

眼底影像跨模态学习医学影像深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。