arXiv:2607.23972cs.CV2026-07

系统梳理眼底彩照AI发展中的数据、预处理与模型协同演进。

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

  • 构建数据-预处理-模型联动视角,揭示三者协同演化路径。
  • 大型多中心数据集融合多模态与纵向病历,支持更全面分析。
  • 适合关注医学影像智能、跨模态建模与临床落地的研究者。

眼底彩照(CFP)是大规模筛查眼科及系统性疾病的主要无创影像手段。现有综述多孤立总结特定任务算法、数据集或预处理技术,缺乏对三者在现代AI背景下的协同发展审视。本文从数据演进、预处理范式与建模框架的互动出发,系统梳理CFP人工智能的发展脉络。数据显示,CFP数据集已从早期单中心小规模标注数据,演变为包含多中心、多模态配对和纵向临床记录的大型资源;预处理由传统图像增强,发展为神经数据工程管线、硬件感知令牌优化及针对不完整电子健康记录(EHRs)的自监督补全;建模则从卷积神经网络(CNNs)迈向视觉基础模型、状态空间模型(SSMs)与多专家架构。在多模态前沿,CFP正日益与EHR和长期患者信息融合,实现超越单一图像分析的综合临床推理。结论指出,未来突破依赖于数据、预处理与多模态建模的协同优化,为实现鲁棒临床部署、跨域泛化提升与高效边缘智能提供路线图。

原文摘要 · Abstract (English)

Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence. This review provides an integrated overview of CFP AI through the interplay of dataset evolution, preprocessing paradigms, and modeling frameworks. We show that CFP datasets have evolved from small single-center collections with task-specific labels to large multi-center resources featuring multimodal pairings and longitudinal clinical records. Preprocessing has progressed from conventional image enhancement to neural data-engineering pipelines, hardware-aware token optimization, and self-supervised imputation for incomplete electronic health records (EHRs). Meanwhile, modeling has advanced from convolutional neural networks (CNNs) to vision foundation models, state space models (SSMs), and multimodal expert architectures. At the multimodal frontier, CFP is increasingly integrated with EHRs and longitudinal patient information, enabling more comprehensive clinical reasoning beyond isolated image analysis. We conclude that future progress depends on the collaborative optimization of datasets, preprocessing, and multimodal modeling, providing a roadmap toward robust clinical deployment, improved cross-domain generalization, and resource-efficient edge intelligence.

医学影像多模态眼底图像AI医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。