arXiv:2409.15087eess.IVcs.CV2024-09被引 14

AI辅助眼病诊断提升准确率与效率,且持续学习让模型更适应不同人群。

AI Workflow, External Validation, and Development in Eye Disease Diagnosis

  • 构建AI辅助诊疗流程,结合真实患者数据验证效果。
  • 24名医生使用AI后平均准确率提升20%,部分案例超50%。
  • 模型通过持续学习在新加坡人群上表现更好,适合跨区域应用。

及时诊断眼病面临疾病负担加重和临床资源不足的挑战。尽管人工智能在提高诊断准确性方面展现出潜力,但其在真实临床流程和多样人群中的有效性仍缺乏充分验证。本研究以老年黄斑变性(AMD)的诊断与严重程度分级为例,设计并实施了AI辅助诊断流程,对比了24名来自12家机构的临床医生在使用与不使用AI辅助时的表现,数据来自年龄相关眼病研究(AREDS)的真实患者样本。同时,通过引入约4万张新增医学图像(命名为AREDS2数据集),对现有AI模型进行持续优化,并在AREDS、AREDS2及新加坡外部测试集上系统评估。结果显示,AI辅助使23名医生诊断准确率显著提升,平均F1分数从手动诊断的37.71升至45.52(P值 < 0.0001),个别案例提升超过50%;17名追踪医生的诊断时间减少,最快节省40%。具备持续学习能力的模型在三个独立数据集上表现稳健,准确率提升29%,新加坡人群中F1分数从42增至54。

原文摘要 · Abstract (English)

Timely disease diagnosis is challenging due to increasing disease burdens and limited clinician availability. AI shows promise in diagnosis accuracy but faces real-world application issues due to insufficient validation in clinical workflows and diverse populations. This study addresses gaps in medical AI downstream accountability through a case study on age-related macular degeneration (AMD) diagnosis and severity classification. We designed and implemented an AI-assisted diagnostic workflow for AMD, comparing diagnostic performance with and without AI assistance among 24 clinicians from 12 institutions with real patient data sampled from the Age-Related Eye Disease Study (AREDS). Additionally, we demonstrated continual enhancement of an existing AI model by incorporating approximately 40,000 additional medical images (named AREDS2 dataset). The improved model was then systematically evaluated using both AREDS and AREDS2 test sets, as well as an external test set from Singapore. AI assistance markedly enhanced diagnostic accuracy and classification for 23 out of 24 clinicians, with the average F1-score increasing by 20% from 37.71 (Manual) to 45.52 (Manual + AI) (P-value < 0.0001), achieving an improvement of over 50% in some cases. In terms of efficiency, AI assistance reduced diagnostic times for 17 out of the 19 clinicians tracked, with time savings of up to 40%. Furthermore, a model equipped with continual learning showed robust performance across three independent datasets, recording a 29% increase in accuracy, and elevating the F1-score from 42 to 54 in the Singapore population.

AI诊断眼科持续学习临床验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。