arXiv:2607.28608cs.LGq-bio.QM2026-07

提出可复现的临床风险模型子群体公平性审计框架,揭示现有方法在不同条件下的可靠性边界。

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

论文配图:KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models
图 1 · 摘自论文原文
  • 构建五阶段审计流程,涵盖分组、差异度量、机制诊断、后处理缓解和漂移监测
  • 发现显著性检验需按各维度最小可检测效应调整,否则易误判公平性问题
  • 强调应报告结果方差而非均值,因校准方法效果随机且不可靠

临床风险模型虽整体表现良好,但在不同患者子群体间误差率差异显著。现有审计流程组件极少经过压力测试,难以判断其可信度与适用条件。本文提出KAISEN,一个五阶段审计流程,包括子群体分层、差异度量、机制诊断、事后缓解和漂移监测,在包含16个疾病任务、15个社会决定因素轴(源自健康人群2030)及三个预设交集的合成基准上评估至失效点。四项发现:(i) 显著性检验需基于各轴自身最小可检测效应进行校准,标准化后的等机会差异(EOD)与显著性计数相关性从rho=0.56升至0.78;(ii) 每组阈值优化在48次独立实验中全部降低EOD(平均差值delta=-0.285,95%置信区间[-0.313, -0.252]),而组间Platt缩放仅在19/48次提升,平均效应接近零,应报告方差而非均值;(iii) 机制诊断在144个受控案例中全正确,但在48个模型驱动案例中因代理变量误设完全失效,无失败信号;(iv) CUSUM的误报与漏报主要受队列实现影响(卡方p=0.002),同一阈值无法跨队列迁移。所有结果基于已知真实值的合成数据,不具临床有效性。代码、工具和脚本均已公开以复现全部结果。

原文摘要 · Abstract (English)

Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Four findings follow. (i) Significance tracks each axis's gap against its own minimum detectable effect: rank correlation between significance count and raw equalized-odds difference (EOD) across the 15 axes is rho = 0.56, rising to rho = 0.78 once EOD is standardized by that floor. (ii) Per-group threshold optimization reduces EOD in 48 of 48 held-out runs (paired delta = -0.285, 95% CI [-0.313, -0.252]), while group-wise Platt scaling -- the better calibrator -- behaves as a coin flip on EOD (19 of 48 runs improved, 95% CI [0.26, 0.55]) with mean effect near zero, so what an audit should report is the variance, not the average. (iii) The mechanism diagnostic classifies 144 of 144 controlled cases correctly but recovers none of 48 model-driven cases under proxy misspecification, with no signal that it failed. (iv) CUSUM failures and false alarms track cohort realization far more than disease: at the reference threshold, all 27 false alarms and 7 of 8 missed shifts come from different seeds (chi-squared p = 0.002), so a threshold tuned on one cohort fails to transfer. All results are synthetic with known ground truth and do not establish clinical validity. Code, artifacts, and scripts reproducing every number are released.

公平性审计临床建模可复现性子群体分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。