arXiv:2608.30846cs.AI2026-08中稿 · the 35th ACM Inter…

提出新方法评估医疗预测公平性审计的结论稳定性。

VFR-Audit: Verdict-Level Reliability for Fairness Audits in Hospital Length-of-Stay Prediction

论文配图:VFR-Audit: Verdict-Level Reliability for Fairness Audits in Hospital Length-of-Stay Prediction
图 1 · 摘自论文原文
  • 用可解释的翻转率衡量审计结论是否可靠
  • 发现传统方法无法保证结论稳定,易随数据波动变化
  • 适合医院管理者、医保机构和监管者做决策参考

临床AI中的公平性审计将连续公平性指标转换为通过/不通过的二元结论,由医院管理委员会、支付方和监管机构据此行动。此类审计需在时间与不同医院间重复进行,导致同一模型在不同审计中可能在通过与不通过之间反复切换。现有不确定性方法如贝叶斯后验、自助法置信区间和置换检验仅能处理连续指标层面的不确定性,无法自动转化为对结论稳定性的判断,且难以在(模型、指标、属性)组合的多维空间中规模化应用。此外,现有方法未解决偏见缓解策略(如重加权或组内阈值调整)是否会以牺牲模型判别力(如AUROC、AUPRC)为代价换取稳定的通过结论的问题。为此,本文提出VFR-Audit框架,核心是判决翻转率(Verdict Flip Rate, VFR),一个取值范围在0到0.5之间的标量,用于衡量在分层自助抽样下判决反转的概率。VFR-Audit同时报告三个可靠性维度:同组内重抽样稳定性、审计规模敏感性,以及跨医院判决一致性(通过Fleiss' kappa评估)。

原文摘要 · Abstract (English)

Fairness audits in clinical Artificial Intelligence convert continuous fairness metrics into binary pass-or-fail verdicts against operational thresholds, where hospital governance boards, payers, and regulators act on the resulting verdicts. Such audits are repeated over time and across hospital sites, thus the same verdict can flip between pass and fail across audits. Existing uncertainty methods such as Bayesian posteriors, bootstrap confidence intervals, and permutation tests address verdict instability only at the continuous-metric level. Converting metric-level uncertainty into a verdict-stability claim remains a manual step that scales poorly across the (model, metric, attribute) cells an audit covers. Existing uncertainty methods also leave open whether bias-mitigation steps, such as reweighing or per-group threshold shifts, yield a stable passing verdict at the cost of model discrimination measured as AUROC or AUPRC.To address this verdict-stability gap, we propose VFR-Audit, a framework built around the Verdict Flip Rate (VFR), a scalar bounded between 0 and 0.5 that measures the probability of verdict reversal under stratified bootstrap resampling. VFR-Audit reports VFR alongside three reliability axes, namely within-cohort resampling stability, audit-size sensitivity, and cross-hospital verdict agreement via Fleiss' kappa.

公平性审计医疗AI可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。