综述眼科多模态AI诊断进展,从专用模型到大模型。
A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models
- 分任务导向与大模型两类,融合眼底彩照、OCT等多模态数据
- 大模型支持跨模态理解与自动生成临床报告,提升诊断效率
- 适合医学AI研究者与临床医生了解技术演进与挑战
视觉障碍是全球重大健康挑战,多模态成像能提供互补信息,对精准眼科诊断至关重要。本文系统综述截至2025年的多模态深度学习在眼科学中的最新进展,涵盖两大类:任务特定的多模态方法与大规模多模态基础模型。前者针对病灶检测、疾病诊断、图像合成等具体临床应用,整合彩色眼底照相、光学相干断层扫描(OCT)和血管造影等多种成像模态;后者结合先进的视觉-语言架构与在多样化眼科数据集上预训练的大语言模型,实现强大的跨模态理解、自动临床报告生成与决策支持。文章还深入分析关键数据集、评估指标及方法创新,包括自监督学习、基于注意力的融合机制与对比对齐策略,并讨论当前挑战:数据异质性、标注有限、可解释性差以及在不同人群间的泛化能力不足。最后展望未来方向,强调超广角成像与基于强化学习的推理框架,以构建智能、可解释且临床可用的AI系统。
原文摘要 · Abstract (English)
Visual impairment represents a major global health challenge, with multimodal imaging providing complementary information that is essential for accurate ophthalmic diagnosis. This comprehensive survey systematically reviews the latest advances in multimodal deep learning methods in ophthalmology up to the year 2025. The review focuses on two main categories: task-specific multimodal approaches and large-scale multimodal foundation models. Task-specific approaches are designed for particular clinical applications such as lesion detection, disease diagnosis, and image synthesis. These methods utilize a variety of imaging modalities including color fundus photography, optical coherence tomography, and angiography. On the other hand, foundation models combine sophisticated vision-language architectures and large language models pretrained on diverse ophthalmic datasets. These models enable robust cross-modal understanding, automated clinical report generation, and decision support. The survey critically examines important datasets, evaluation metrics, and methodological innovations including self-supervised learning, attention-based fusion, and contrastive alignment. It also discusses ongoing challenges such as variability in data, limited annotations, lack of interpretability, and issues with generalizability across different patient populations. Finally, the survey outlines promising future directions that emphasize the use of ultra-widefield imaging and reinforcement learning-based reasoning frameworks to create intelligent, interpretable, and clinically applicable AI systems for ophthalmology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。