arXiv:2506.17035cs.LG2025-06被引 8

梳理临床AI公平性度量,揭示现有方法的局限与空白。

Critical Appraisal of Fairness Metrics in Clinical Predictive AI

  • 系统整理62个公平性度量,按性能依赖等分类
  • 仅1个度量关注临床实用性,多数缺乏真实验证
  • 指出需加强不确定性、交叉歧视和实际应用研究

预测性人工智能为改善临床实践和患者结果带来机遇,但若公平性未被妥善处理,可能加剧偏见。然而,“公平性”的定义仍不明确。我们开展了一项范围综述,旨在识别并批判性评估临床预测型AI中的公平性度量。我们将“公平性度量”定义为量化模型是否对基于敏感属性(如性别、种族)的个体或群体存在社会性歧视的指标。通过检索五个数据库(2014–2024),筛选820条记录,最终纳入41项研究,提取出62个公平性度量。这些度量依据性能依赖性、模型输出层级及基础性能指标进行分类,揭示出该领域碎片化严重,临床验证不足,且过度依赖阈值相关度量。其中18个度量专为医疗场景设计,但仅有1个涉及临床实用性。研究结果凸显了公平性定义与量化中的概念挑战,并指出在不确定性量化、交叉歧视及现实适用性方面存在显著缺口。未来工作应优先发展具有临床意义的度量。

原文摘要 · Abstract (English)

Predictive artificial intelligence (AI) offers an opportunity to improve clinical practice and patient outcomes, but risks perpetuating biases if fairness is inadequately addressed. However, the definition of "fairness" remains unclear. We conducted a scoping review to identify and critically appraise fairness metrics for clinical predictive AI. We defined a "fairness metric" as a measure quantifying whether a model discriminates (societally) against individuals or groups defined by sensitive attributes. We searched five databases (2014-2024), screening 820 records, to include 41 studies, and extracted 62 fairness metrics. Metrics were classified by performance-dependency, model output level, and base performance metric, revealing a fragmented landscape with limited clinical validation and overreliance on threshold-dependent measures. Eighteen metrics were explicitly developed for healthcare, including only one clinical utility metric. Our findings highlight conceptual challenges in defining and quantifying fairness and identify gaps in uncertainty quantification, intersectionality, and real-world applicability. Future work should prioritise clinically meaningful metrics.

AI公平性临床预测度量评估医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。