用眼底照片生成3D眼底断层图,提升眼科诊疗可及性
APTOS-2024 challenge report: Generation of synthetic 3D OCT images from fundus photographs
- 通过跨模态协同与视觉基础模型,将2D眼底图转为3D OCT影像
- 提出像素级与语义级双评估指标,量化生成图像真实性
- 首次验证该技术在资源匮乏地区应用的可行性,适合医疗研发者
光学相干断层扫描(OCT)可无创获取高分辨率、三维视网膜结构影像,是病灶定位与疾病诊断的关键工具,但受限于设备成本和专业操作需求。相比之下,二维彩色眼底照相具有采集快、易普及的优势。尽管生成式人工智能在医学图像合成中表现优异,但将2D眼底图像转换为3D OCT图像仍面临模态间维度与生物信息差异的挑战。为此,亚太远程眼科协会(APTOS-2024)组织了‘基于AI的眼底图像生成OCT’挑战赛。本文详述了挑战赛框架(称作APTOS-2024 Challenge),包括基准数据集、两项保真度评估指标——基于图像的距离(像素级OCT B-scan相似性)和基于视频的距离(语义级体积一致性),以及顶尖方案分析。共吸引342支团队参与,提交42份初赛作品,9支进入决赛。领先方法融合了混合数据预处理/增强(跨模态协同范式)、外部眼科影像数据集预训练、视觉基础模型整合及模型架构优化。APTOS-2024 Challenge是首个证明眼底图到3D OCT合成可行性的基准,有望提升欠发达地区眼科服务可及性,并加速医学研究与临床应用。
原文摘要 · Abstract (English)
Optical Coherence Tomography (OCT) provides high-resolution, 3D, and non-invasive visualization of retinal layers in vivo, serving as a critical tool for lesion localization and disease diagnosis. However, its widespread adoption is limited by equipment costs and the need for specialized operators. In comparison, 2D color fundus photography offers faster acquisition and greater accessibility with less dependence on expensive devices. Although generative artificial intelligence has demonstrated promising results in medical image synthesis, translating 2D fundus images into 3D OCT images presents unique challenges due to inherent differences in data dimensionality and biological information between modalities. To advance generative models in the fundus-to-3D-OCT setting, the Asia Pacific Tele-Ophthalmology Society (APTOS-2024) organized a challenge titled Artificial Intelligence-based OCT Generation from Fundus Images. This paper details the challenge framework (referred to as APTOS-2024 Challenge), including: the benchmark dataset, evaluation methodology featuring two fidelity metrics-image-based distance (pixel-level OCT B-scan similarity) and video-based distance (semantic-level volumetric consistency), and analysis of top-performing solutions. The challenge attracted 342 participating teams, with 42 preliminary submissions and 9 finalists. Leading methodologies incorporated innovations in hybrid data preprocessing or augmentation (cross-modality collaborative paradigms), pre-training on external ophthalmic imaging datasets, integration of vision foundation models, and model architecture improvement. The APTOS-2024 Challenge is the first benchmark demonstrating the feasibility of fundus-to-3D-OCT synthesis as a potential solution for improving ophthalmic care accessibility in under-resourced healthcare settings, while helping to expedite medical research and clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。