arXiv:2608.30568cs.LGcs.AI2026-08

AUC等指标的非可合并性可能误导公平性评估,需警惕整体性能与子群表现的偏差。

Collapsibility of Performance Metrics in Clinical Predictive AI

论文配图:Collapsibility of Performance Metrics in Clinical Predictive AI
图 1 · 摘自论文原文
  • 识别出5个非可合并指标,其中AUC因跨组效应导致整体值超范围
  • 发现整体AUC可能高于或低于所有子群AUC,违背直觉
  • 提醒临床AI开发者关注指标特性,避免误判模型公平性

人群层面的预测人工智能(AI)性能评估可能掩盖子群间的差异。公平性评价通常依赖于子群间的表现分析。然而,部分性能指标具有不可合并性,即总体性能值不等于各子群特定值的加权平均。本文研究了15种常用性能指标的可合并性,通过线性分解或基于辛普森悖论的反例进行形式化证明。结果显示,5个指标(AUC、校准截距、校准斜率、期望校准误差、Nagelkerke R²)为非可合并,其余10个(O:E比、logloss、Brier分数、准确率、F1分数、真阳性率、真阴性率、阳性预测值、阴性预测值、净收益)为可合并。特别地,由于当子群体共存时AUC可分解为组内与组间两部分,其总体值可能超出各子群值的范围。结论:性能指标的不可合并性对报告、模型评估和公平性审查有重要影响,可能导致子群与总体性能间的虚假差异,从而误导公平性判断。明确承认并报告指标的可合并性特征,有助于提升公平性评估的可解释性和透明度。

原文摘要 · Abstract (English)

Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI, with a focus on the area under the receiver operating characteristic curve (AUC, also known as c-statistic). Methods: We investigate the collapsibility of 15 performance metrics, either by expressing each metric as a linear combination of its stratum specific values or, where non-collapsible, by providing a counterexample inspired by Simpson's paradox as a formal disproof. Results: Five performance metrics (AUC, calibration intercept, calibration slope, expected calibration error, and Nagelkerke R^2) are shown to be non-collapsible, and ten (O:E ratio, logloss, Brier score, accuracy, F1-score, true positive rate, true negative rate, positive predictive value, negative predictive value, and net benefit) are shown to be collapsible. The AUC is shown to be non-collapsible because it decomposes into within- and cross-group AUC terms when subpopulations coexist, such that its overall value may fall outside the range of subgroup specific AUCs. Conclusions: Non-collapsibility of performance metrics has important consequences for reporting, model appraisal, and fairness evaluation. It can generate spurious differences between subgroup and overall performance, which may mislead fairness evaluations. Explicitly acknowledging and reporting the collapsibility properties of performance metrics improves both the interpretability and transparency of fairness assessments.

AI公平性性能评估医学AIAUC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。