AI模型比医生更准判断卵巢肿块良恶性,人机协作效果最佳。
From ACR O-RADS 2022 to Explainable Deep Learning: Comparative Performance of Expert Radiologists, Convolutional Neural Networks, Vision Transformers, and Fusion Models in Ovarian Masses
- 用深度学习和视觉变压器分析超声图像,自动分类卵巢肿块。
- ViT模型准确率达87.4%,远超医生单独判断的68.0%。
- 医生评分与AI结合后诊断效果最优,适合临床辅助决策。
背景:2022版卵巢-附件报告与数据系统(O-RADS)更新了对附件病变的风险分层,但人工解读仍存在变异性及保守阈值问题。与此同时,深度学习(DL)模型在图像引导的卵巢病变表征中展现出潜力。本研究评估了应用O-RADS v2022的放射科医生表现,对比了领先的卷积神经网络(CNN)与视觉变换器(ViT)模型,并探究了混合人机框架的诊断增益。方法:单中心回顾性队列研究,纳入227例患者共512张卵巢肿块超声图像(110例至少有一个恶性囊肿)。训练并验证了16种DL模型,包括DenseNets、EfficientNets、ResNets、VGGs、Xception和ViTs。还为每种方案构建了融合放射科医生O-RADS评分与DL预测概率的混合模型。结果:仅由放射科医生进行的O-RADS评估达到AUC 0.683,总体准确率68.0%。CNN模型的AUC范围为0.620至0.908,准确率59.2%至86.4%;其中ViT16-384表现最佳,AUC达0.941,准确率为87.4%。混合人机框架显著提升了CNN模型性能,但对ViT模型的提升未达统计显著性(P>0.05)。结论:深度学习模型显著优于仅靠医生的O-RADS v2022评估,将专家评分与人工智能结合可获得最高诊断准确率与判别力。混合人机范式有望标准化盆腔超声解读,减少假阳性,提升高风险病变检出能力。
原文摘要 · Abstract (English)
Background: The 2022 update of the Ovarian-Adnexal Reporting and Data System (O-RADS) ultrasound classification refines risk stratification for adnexal lesions, yet human interpretation remains subject to variability and conservative thresholds. Concurrently, deep learning (DL) models have demonstrated promise in image-based ovarian lesion characterization. This study evaluates radiologist performance applying O-RADS v2022, compares it to leading convolutional neural network (CNN) and Vision Transformer (ViT) models, and investigates the diagnostic gains achieved by hybrid human-AI frameworks. Methods: In this single-center, retrospective cohort study, a total of 512 adnexal mass images from 227 patients (110 with at least one malignant cyst) were included. Sixteen DL models, including DenseNets, EfficientNets, ResNets, VGGs, Xception, and ViTs, were trained and validated. A hybrid model integrating radiologist O-RADS scores with DL-predicted probabilities was also built for each scheme. Results: Radiologist-only O-RADS assessment achieved an AUC of 0.683 and an overall accuracy of 68.0%. CNN models yielded AUCs of 0.620 to 0.908 and accuracies of 59.2% to 86.4%, while ViT16-384 reached the best performance, with an AUC of 0.941 and an accuracy of 87.4%. Hybrid human-AI frameworks further significantly enhanced the performance of CNN models; however, the improvement for ViT models was not statistically significant (P-value >0.05). Conclusions: DL models markedly outperform radiologist-only O-RADS v2022 assessment, and the integration of expert scores with AI yields the highest diagnostic accuracy and discrimination. Hybrid human-AI paradigms hold substantial potential to standardize pelvic ultrasound interpretation, reduce false positives, and improve detection of high-risk lesions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。