arXiv:2605.02714cs.CVcs.AI2026-05被引 1

OphMAE融合三维与二维眼底影像,实现自适应眼科诊断

OphMAE: Bridging Volumetric and Planar Imaging with a Foundation Model for Adaptive Ophthalmological Diagnosis

论文配图:OphMAE: Bridging Volumetric and Planar Imaging with a Foundation Model for Adaptive Ophthalmological Diagnosis
图 1 · 摘自论文原文
  • 设计跨模态融合架构,统一处理3D OCT与2D眼底图像
  • 在17项任务中达到96.9%(AMD)和97.2%(DME)AUC
  • 仅用500样本即可保持95.7% AUC,适合资源受限场景

基础模型的兴起开启了医疗AI新纪元,使大规模无标签数据中提取可泛化表征成为可能。然而,现有眼科AI多局限于单一模态推理,与临床需整合多种影像的实际需求不符。此外,高性能AI在资源匮乏地区常因缺乏先进三维成像设备而难以部署。本文提出眼科多模态掩码自编码器(OphMAE),通过创新的跨模态融合架构与自适应推理机制,将3D光学相干断层扫描(OCT)的深度信息与2D眼底平面图像的上下文结合。模型在包含32,765名患者、共183,875对OCT图像的大规模数据集上预训练。在涵盖17个诊断任务、48,340对图像、8,191名患者的严格基准测试中,其对年龄相关性黄斑变性(AMD)的曲线下面积(AUC)达96.9%,对糖尿病性黄斑水肿(DME)达97.2%,显著优于现有单模态与多模态基础模型。关键优势在于工程适应性强:即使仅输入2D图像,对AMD的AUC仍达93.7%;且具备极强数据效率,仅需500个标注样本即保持95.7% AUC。该工作建立了一个可扩展、自适应的眼科AI框架,确保在不同任务中均保持高鲁棒性。

原文摘要 · Abstract (English)

The advent of foundation models has heralded a new era in medical artificial intelligence (AI), enabling the extraction of generalizable representations from large-scale unlabeled datasets. However, current ophthalmic AI paradigms are predominantly constrained to single-modality inference, thereby creating a dissonance with clinical practice where diagnosis relies on the synthesis of complementary imaging modalities. Furthermore, the deployment of high-performance AI in resource-limited settings is frequently impeded by the unavailability of advanced three-dimensional imaging hardware. Here, we present the Ophthalmic multimodal Masked Autoencoder (OphMAE), a multi-imaging foundation model engineered to synergize the volumetric depth of 3D Optical Coherence Tomography (OCT) with the planar context of 2D en face OCT. By implementing a novel cross-modal fusion architecture and a unique adaptive inference mechanism, OphMAE was pre-trained on a massive dataset with of 183,875 paired OCT images derived from 32,765 patients. In a rigorous benchmark encompassing 17 diverse diagnostic tasks with 48,340 paired OCT images from 8,191 patients, the model demonstrated state-of-the-art performance, achieving an Area Under the Curve (AUC) of 96.9% for Age-related Macular Degeneration (AMD) and 97.2% for Diabetic Macular Edema (DME), consistently surpassing existing single-modal and multimodal foundation models. Crucially, OphMAE exhibits robust engineering adaptability: it maintains high diagnostic accuracy, such as 93.7\% AUC for AMD, even when restricted to single-modality 2D inputs, and demonstrates exceptional data efficiency by retaining 95.7% AUC with as few as 500 labeled samples. This work establishes a scalable and adaptable framework for ophthalmic AI, ensuring robust performance across different tasks.

眼科AI多模态自编码器数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。