arXiv:2606.15019cs.CV2026-06

首个跨国验证的宫颈癌筛查深度学习模型,提升基层医疗诊断能力。

Towards Global AI-Driven Cervical Cancer Screening

论文配图:Towards Global AI-Driven Cervical Cancer Screening
图 1 · 摘自论文原文
  • 将病变检测与分类建模为多任务学习,联合图像分类与病灶分割。
  • 在本国数据上超越医生准确率(0.68 vs 0.64),跨国家验证表现稳定。
  • 适合资源匮乏地区使用,尤其关注合并症对模型性能的显著影响。

全球消除宫颈癌是世界卫生组织(WHO)设定的关键公共卫生目标,筛查可使死亡率降低多达80%。然而,在中低收入国家,专业医生和活检服务资源有限。基于深度学习(DL)的算法为筛查提供了有前景的支持,但现有方法大多仅在单一国家的私有数据集上开发和验证。本文首次提出并验证了在多国数据上适用的宫颈癌筛查深度学习方法。技术上,将宫颈镜图像中的病变检测与分类问题建模为多任务学习,同时完成图像级分类与病变分割。模型在酸染色宫颈镜图像的私有数据集上训练,采用人工标注的病灶掩码及相应病理结果,并通过大规模数据增强应对图像差异。在以病理结果为真实标签的内部验证中,算法在区分CIN1-与CIN2+的分类任务上表现优于医学专家(平衡准确率:0.68 vs 0.64)。在四个来自不同国家、流行率与患者特征差异显著的外部数据集上进行验证,本方法整体性能优于基线模型。各国间性能差异较大,AUC值范围为0.54至0.80。总体而言,模型表现受年龄、转化区状态、共病及典型征象影响,其中共病的影响最为显著。未来工作应聚焦于提升模型鲁棒性与泛化能力。

原文摘要 · Abstract (English)

The global elimination of cervical cancer is a key public health goal set by the World Health Organization (WHO), with screening programs reducing mortality by up to 80%. However, access to experts and biopsy services is limited in low- to middle-income countries (LMICs). Deep learning (DL)-based algorithms offer promising support for screening, but most existing approaches have been developed and validated on private datasets from single countries. We present the first DL-based approach to cervical cancer screening validated on data from multiple countries. Technically, we phrase the problem of detecting and classifying lesions in colposcopy images as a multi-task learning problem, in which we simultaneously perform image-level classification and lesion segmentation. Our model was trained on a private data set of acid stain colposcopy images with manually generated lesion segmentation masks and corresponding histopathological results, employing extensive data augmentation to address image variability. In an in-distribution validation with pathology results serving as ground truth, our algorithm outperformed medical experts (Balanced Accuracy: 0.68 vs 0.64) in CIN1- (Cervical intraepithelial neoplasia grade 1 or lower) versus CIN2+ (grade 2 or higher) classification. External validation on four colposcopy data sets from four countries featuring radical differences in prevalence and patient characteristics yielded superior performance of our method compared to baseline methods. Performance variability across countries was high with AUC values ranging from 0.54 - 0.80. Overall, algorithm performance varied with age, transformation zone (cervical area most prone to lesion development), presence of comorbidities and pathognomonic signs, with comorbidities having by far the largest negative effect. Future work should focus on improving model robustness and generalizability.

宫颈癌筛查多国验证深度学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。