系统梳理视网膜OCT图像表示学习方法,助力自动诊断与跨设备一致性
Representation learning from OCT images

- 按学习范式分类综述从传统CNN到大模型的多种方法
- 涵盖自监督、生成式、多模态等主流技术路径
- 适合医学影像、AI医疗研究者参考,关注公平性与可解释性
光学相干断层扫描(OCT)已成为眼科最常用的成像方式,能提供高分辨率、非侵入性的视网膜微结构可视化。通过表示学习实现OCT图像的自动化分析已成为研究前沿,主要源于临床对处理大规模数据的需求。目标是减少对专家标注的依赖,并提升不同设备和人群间的诊断一致性。本文系统回顾了视网膜OCT图像分析中表示学习方法的发展,涵盖从早期深度学习到最新基础模型与视觉-语言系统。按学习范式组织文献,包括基于CNN和Transformer的监督学习、自监督与半监督方法、生成式方法、3D体建模、多模态表示学习及大规模预训练基础模型。每类分析核心贡献,识别持续局限,并追踪各方法间的演进关系。此外,梳理公开OCT数据集,讨论评估协议,提出统一问题形式化框架。基于此分析,指出当前关键开放方向:体素级基础模型预训练、不确定性感知表示学习、联邦与隐私保护训练、公平性与偏见缓解、概念可解释性等。
原文摘要 · Abstract (English)
Optical Coherence Tomography (OCT) has become one of the most used imaging modality in ophthalmology. It provides high-resolution, non-invasive visualization of retinal microarchitecture. The automated analysis of OCT images through representation learning has emerged as a central research frontier. This has mainly been driven by the clinical need to process large acquisition volumes. The objective is to reduce the reliance on expert annotation, and improve diagnostic consistency across devices and populations. This survey provides a comprehensive and structured review of representation learning methods for retinal OCT image analysis. It covers the period from early deep learning approaches to the most recent developments in foundation models and vision-language systems. We organize the literature along a principled taxonomy of learning paradigms, encompassing supervised learning with CNN-based and transformer-based architectures, self-supervised and semi-supervised methods, generative approaches, as well as 3D volumetric modeling, multimodal representation learning, and large-scale pretrained foundation models. For each paradigm, we analyze the core methodological contributions, identify persistent limitations, and trace the connections between successive approaches. We further provide a structured overview of publicly available OCT datasets, discuss evaluation protocol considerations, and present a unified problem formulation that situates each learning paradigm within a common mathematical framework. Building on this analysis, we identify and discuss the most pressing open research directions emerging in the literature. This includes volumetric foundation model pretraining, uncertainty-aware representation learning, federated and privacy-preserving training, fairness and bias mitigation, concept-based interpretability,...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。