AI比传统方法更可靠地检测病理切片组织,避免漏检。
The impact of tissue detection on diagnostic artificial intelligence algorithms in digital pathology
- 用AI和传统阈值法分别做组织检测,对比其对诊断模型的影响。
- AI检测使完全漏检的切片从0.43%降至0.08%,显著减少失败率。
- 虽整体评分相似,但3.5%恶性切片因检测差异出现临床重要偏差。
组织检测是数字病理学中大多数应用的关键第一步。现有研究极少报告分割算法细节,也缺乏对劣质分割算法下游影响的探讨。忽略组织检测质量可能成为下游性能瓶颈,在临床应用中若遗漏诊断相关区域,将危及患者安全。本研究旨在评估下游任务对组织检测方法的敏感性,并比较经典方法与AI方法的表现。我们基于5台扫描仪采集的33,823张全切片图像(WSIs)训练了用于前列腺癌格里森分级的AI模型,采用两种组织检测算法:阈值法(经典)和UNet++(AI)。下游格里森分级模型在13个临床中心、13种扫描仪采集的70,524张WSI上进行训练与测试。从阈值法切换至AI方法后,完全未检测到组织的样本从116例(0.43%)降至22例(0.08%),表明AI模型在异常外观切片上更具鲁棒性。在两种方法均能检测到组织的切片中,整体格里森分级性能无显著差异。然而,在3.5%的恶性切片中,检测方式导致了具有临床意义的评级差异,凸显稳健组织检测对诊断AI最优临床表现的重要性。
原文摘要 · Abstract (English)
Tissue detection is a crucial first step in most digital pathology applications. Details of the segmentation algorithm are rarely reported, and there is a lack of studies investigating the downstream effects of a poor segmentation algorithm. Disregarding tissue detection quality could create a bottleneck for downstream performance and jeopardize patient safety if diagnostically relevant parts of the specimen are excluded from analysis in clinical applications. This study aims to determine whether performance of downstream tasks is sensitive to the tissue detection method, and to compare performance of classical and AI-based tissue detection. To this end, we trained an AI model for Gleason grading of prostate cancer in whole slide images (WSIs) using two different tissue detection algorithms: thresholding (classical) and UNet++ (AI). A total of 33,823 WSIs scanned on five digital pathology scanners were used to train the tissue detection AI model. The downstream Gleason grading algorithm was trained and tested using 70,524 WSIs from 13 clinical sites scanned on 13 different scanners. There was a decrease from 116 (0.43%) to 22 (0.08%) fully undetected tissue samples when switching from thresholding-based tissue detection to AI-based, suggesting an AI model may be more reliable than a classical model for avoiding total failures on slides with unusual appearance. On the slides where tissue could be detected by both algorithms, no significant difference in overall Gleason grading performance was observed. However, tissue detection dependent clinically significant variations in AI grading were observed in 3.5% of malignant slides, highlighting the importance of robust tissue detection for optimal clinical performance of diagnostic AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。