提出新评估方法,让癌症筛查AI在不同人群间更公平可靠
Cohort-attention Evaluation Metrics for Tied Data
- 按患者层面评估,结合熵加权分布
- 提升跨群体敏感性与特异性,避免数据偏差
- 适合医疗AI模型评估,尤其关注公平性
人工智能显著提升了癌症检测与风险评估的准确率。然而,传统分类指标常忽视数据不平衡、群体间表现差异及患者层面不一致问题,导致评估偏差。本文提出针对关联数据的队列注意力评估指标(CAT),引入患者级评估、基于熵的分布加权,以及队列加权的敏感性和特异性。关键指标如CAT敏感性、CAT特异性与CAT均值,确保在多样化人群中的均衡与公正评估。该方法增强了预测的可靠性、公平性与可解释性,为医疗筛查AI模型提供稳健的评估方案。
原文摘要 · Abstract (English)
Artificial intelligence (AI) has significantly improved medical screening accuracy, particularly in cancer detection and risk assessment. However, traditional classification metrics often fail to account for imbalanced data, varying performance across cohorts, and patient-level inconsistencies, leading to biased evaluations. We propose the cohort-attention evaluation metrics for tied data (CAT). CAT introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity. Key metrics like CAT Sensitivity, CAT Specificity, and CAT Mean ensure balanced and fair evaluation across diverse populations. This approach enhances predictive reliability, fairness, and interpretability, providing a robust evaluation method for AI-driven medical screening models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。