arXiv:2512.22242cs.LGcs.AI2025-12中稿 · publication at the…被引 2

评估肺癌筛查AI模型在不同人群中的公平性,发现性别和种族差异导致检测效果不一。

Fairness Evaluation of Risk Estimation Models for Lung Cancer Screening

  • 基于JustEFAB框架,对比多个模型在不同人口群体的表现差异。
  • 女性患者中Sybil模型的AUROC达0.88,男性仅为0.81(p<0.001)。
  • 黑人患者在Venkadesh21模型下敏感度仅0.39,显著低于白人(0.69)

肺癌是全球成人癌症死亡的首要原因。对高危人群每年进行低剂量CT(LDCT)筛查可实现早期发现并降低死亡率,但可能加剧本已紧张的放射科人力负担。人工智能模型在从LDCT影像估算肺癌风险方面展现出潜力,但高危人群多样,模型在不同人口群体中的表现仍存疑问。本研究依据JustEFAB框架中关于混杂因素与伦理重要偏差的考量,评估了两个深度学习风险预测模型——Sybil肺癌风险模型与Venkadesh21结节风险评估器——以及英国胸科学会指南推荐的PanCan2b逻辑回归模型在肺癌筛查中的表现差异。两个深度学习模型均基于美国国家肺筛查试验(NLST)数据训练,并在保留的NLST验证集上评估。我们分析了各人口子群的AUROC、敏感性和特异性,并探讨临床风险因素的潜在混杂效应。结果显示,Sybil模型在女性(AUROC=0.88,95%CI: 0.86, 0.90)与男性(AUROC=0.81,95%CI: 0.78, 0.84)之间存在统计学显著差异(p<0.001)。在90%特异性条件下,Venkadesh21对黑人参与者(敏感度=0.39,95%CI: 0.23, 0.59)的检出率显著低于白人(敏感度=0.69,95%CI: 0.65, 0.73)。这些差异无法由现有临床混杂因素解释,根据JustEFAB标准可视为不公平偏差。研究强调需加强对代表性不足群体的模型性能改进与持续监测,并推动算法公平性研究在肺癌筛查中的应用。

原文摘要 · Abstract (English)

Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95% CI: 0.86, 0.90) and men (0.81, 95% CI: 0.78, 0.84, p < .001). At 90% specificity, Venkadesh21 showed lower sensitivity for Black (0.39, 95% CI: 0.23, 0.59) than White participants (0.69, 95% CI: 0.65, 0.73). These differences were not explained by available clinical confounders and thus may be classified as unfair biases according to JustEFAB. Our findings highlight the importance of improving and monitoring model performance across underrepresented subgroups, and further research on algorithmic fairness, in lung cancer screening.

AI医疗公平性肺癌筛查深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。