arXiv:2506.08900cs.CV2025-06中稿 · publication in npj…被引 25

构建多模态眼底OCT与SLO分析基础模型,提升临床适用性。

MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis

  • 融合OCT与SLO图像的多模态训练,增强模型泛化能力
  • 在分类与分割任务上优于通用及专用模型
  • 提供公开基准与模型,助力眼科AI研发

人工智能已成为辅助临床分析眼底光学相干断层扫描(OCT)图像的重要工具。然而,现有AI模型常需大量标注数据,且在未见数据上表现不佳。基础模型(FMs)通过大规模无标签数据预训练,有望克服这些限制。但当前眼科领域基础模型缺乏充分验证,尤其在分割任务上,且多局限于单一成像模态。为此,我们提出MIRAGE——一种面向OCT与扫描激光检眼镜(SLO)图像的新型多模态基础模型,并构建包含分类与分割任务的新评估基准。与通用及专用基础模型和分割方法对比表明,MIRAGE在两类任务中均表现更优,凸显其作为稳健眼底OCT分析AI系统基础的潜力。MIRAGE与评估基准均已开源:https://github.com/j-morano/MIRAGE。

原文摘要 · Abstract (English)

Artificial intelligence (AI) has become a fundamental tool for assisting clinicians in analyzing ophthalmic images, such as optical coherence tomography (OCT). However, developing AI models often requires extensive annotation, and existing models tend to underperform on independent, unseen data. Foundation models (FMs), large AI models trained on vast unlabeled datasets, have shown promise in overcoming these challenges. Nonetheless, available FMs for ophthalmology lack extensive validation, especially for segmentation tasks, and focus on a single imaging modality. In this context, we propose MIRAGE, a novel multimodal FM for the analysis of OCT and scanning laser ophthalmoscopy (SLO) images. Additionally, we propose a new evaluation benchmark with OCT/SLO classification and segmentation tasks. The comparison with general and specialized FMs and segmentation methods shows the superiority of MIRAGE in both types of tasks, highlighting its suitability as a basis for the development of robust AI systems for retinal OCT image analysis. Both MIRAGE and the evaluation benchmark are publicly available: https://github.com/j-morano/MIRAGE.

眼底图像多模态基础模型OCT分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。