arXiv:2511.11636cs.LGcs.CY2025-11被引 2

开发可解释公平的多囊卵巢综合征风险评估工具,支持临床实时决策。

An Explainable and Fair AI Tool for PCOS Risk Assessment: Calibration, Subgroup Equity, and Interactive Clinical Deployment

  • 结合SHAP与人口统计审计,实现预测结果与诊断差异的关联分析。
  • 随机森林模型准确率达90.8%,在25-35岁女性中表现最佳(90.9%)。
  • 提供交互式网页界面,支持临床场景下的即时风险评估与假设分析。

本文提出一种经过公平性审计且可解释的机器学习框架,用于预测多囊卵巢综合征(PCOS),旨在评估模型性能并识别患者亚组间的诊断差异。该框架融合基于SHAP的特征重要性分析与人口统计学审计,将预测解释与实际差异联系起来,以提供可行动的洞察。采用贝叶斯评分(Brier Score)和期望校准误差(ECE)等概率校准指标,确保各亚组风险预测的可靠性。使用随机森林、SVM和XGBoost模型,通过等距校准(isotonic scaling)和Platt缩放进行校准与公平性对比。校准后的随机森林模型达到90.8%的高预测准确率;SHAP分析显示卵泡数量、体重增加和月经不规律为最关键特征,符合罗切斯特诊断标准。尽管带等距校准的SVM取得最低校准误差(ECE=0.0541),但随机森林在校准与可解释性之间取得更好平衡(Brier=0.0678,ECE=0.0666),因此被选为后续分析模型。亚组分析表明,模型在25-35岁女性中表现最佳(准确率90.9%),但在25岁以下女性中表现较差(准确率69.2%),凸显年龄相关差异。模型在肥胖女性中实现完美精确度,在瘦型PCOS病例中保持高召回率,表现出对不同表型的鲁棒性。最后,基于Streamlit的网页界面实现了实时PCOS风险评估、罗切斯特标准判断及交互式‘如果…会怎样’分析,弥合了人工智能研究与临床应用之间的鸿沟。

原文摘要 · Abstract (English)

This paper presents a fairness-audited and interpretable machine learning framework for predicting polycystic ovary syndrome (PCOS), designed to evaluate model performance and identify diagnostic disparities across patient subgroups. The framework integrated SHAP-based feature attributions with demographic audits to connect predictive explanations with observed disparities for actionable insights. Probabilistic calibration metrics (Brier Score and Expected Calibration Error) are incorporated to ensure reliable risk predictions across subgroups. Random Forest, SVM, and XGBoost models were trained with isotonic and Platt scaling for calibration and fairness comparison. A calibrated Random Forest achieved a high predictive accuracy of 90.8%. SHAP analysis identified follicle count, weight gain, and menstrual irregularity as the most influential features, which are consistent with the Rotterdam diagnostic criteria. Although the SVM with isotonic calibration achieved the lowest calibration error (ECE = 0.0541), the Random Forest model provided a better balance between calibration and interpretability (Brier = 0.0678, ECE = 0.0666). Therefore, it was selected for detailed fairness and SHAP analyses. Subgroup analysis revealed that the model performed best among women aged 25-35 (accuracy 90.9%) but underperformed in those under 25 (69.2%), highlighting age-related disparities. The model achieved perfect precision in obese women and maintained high recall in lean PCOS cases, demonstrating robustness across phenotypes. Finally, a Streamlit-based web interface enables real-time PCOS risk assessment, Rotterdam criteria evaluation, and interactive 'what-if' analysis, bridging the gap between AI research and clinical usability.

可解释AI医疗诊断公平性评估临床部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。