arXiv:2602.03302cs.CVcs.AI2026-02

用大模型实现3D眼底OCT全流程自动诊断,准确率超专家。

Full end-to-end diagnostic workflow automation of 3D OCT via foundation model-driven AI for retinal diseases

  • 基于视觉大模型的统一框架,逐级完成质量评估、病变检测和多病种分类。
  • 在4万张切片上测试,患者级诊断F1达94.39%,跨中心验证稳定在90%以上。
  • 性能媲美专家且更高效,适合大规模眼病筛查应用。

光学相干断层扫描(OCT)凭借高分辨率和三维成像能力革新了视网膜疾病诊断,但临床中仍受限于多阶段流程和传统单切片单任务AI模型。本文提出全流程OCT临床应用系统(FOCUS),一种基于基础模型的端到端框架,可自动化完成3D OCT视网膜疾病诊断。FOCUS依次使用EfficientNetV2-S进行图像质量评估,再通过微调的视觉基础模型实现异常检测与多疾病分类。关键在于采用统一自适应聚合方法,智能融合二维切片级预测,生成完整的三维患者级诊断。在3,300名患者(40,672张切片)上训练并测试,并在四个不同层级中心及多种OCT设备上外部验证1,345名患者(18,498张切片),FOCUS在质量评估(F1: 99.01%)、异常检测(F1: 97.46%)和患者级诊断(F1: 94.39%)上均表现优异。真实世界跨中心验证显示性能稳定(F1: 90.22%-95.24%)。与人工对比中,其异常检测(F1: 95.47% vs 90.91%)和多病种诊断(F1: 93.49% vs 91.35%)达到专家水平,同时效率更高。FOCUS实现了从图像到诊断的全流程自动化,为无人化眼科诊疗提供可验证范本,显著提升大规模人群眼病筛查的可及性与效率。

原文摘要 · Abstract (English)

Optical coherence tomography (OCT) has revolutionized retinal disease diagnosis with its high-resolution and three-dimensional imaging nature, yet its full diagnostic automation in clinical practices remains constrained by multi-stage workflows and conventional single-slice single-task AI models. We present Full-process OCT-based Clinical Utility System (FOCUS), a foundation model-driven framework enabling end-to-end automation of 3D OCT retinal disease diagnosis. FOCUS sequentially performs image quality assessment with EfficientNetV2-S, followed by abnormality detection and multi-disease classification using a fine-tuned Vision Foundation Model. Crucially, FOCUS leverages a unified adaptive aggregation method to intelligently integrate 2D slices-level predictions into comprehensive 3D patient-level diagnosis. Trained and tested on 3,300 patients (40,672 slices), and externally validated on 1,345 patients (18,498 slices) across four different-tier centers and diverse OCT devices, FOCUS achieved high F1 scores for quality assessment (99.01%), abnormally detection (97.46%), and patient-level diagnosis (94.39%). Real-world validation across centers also showed stable performance (F1: 90.22%-95.24%). In human-machine comparisons, FOCUS matched expert performance in abnormality detection (F1: 95.47% vs 90.91%) and multi-disease diagnosis (F1: 93.49% vs 91.35%), while demonstrating better efficiency. FOCUS automates the image-to-diagnosis pipeline, representing a critical advance towards unmanned ophthalmology with a validated blueprint for autonomous screening to enhance population scale retinal care accessibility and efficiency.

OCT诊断大模型自动化眼病筛查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。