arXiv:2605.12895cs.LGcs.AI2026-05

提出RISED框架,用五维评估高风险AI医疗系统,发现传统指标忽略的可靠性与公平性缺陷。

RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare

  • 构建五维度评估框架,通过置信区间与统计校正实现可量化判断
  • 在糖尿病等数据集上发现模型对阈值敏感、群体差异显著等问题
  • 开源工具包支持临床部署前验证,适合医疗AI研发与监管者使用

临床决策支持系统通常仅以单一测试集准确率通过审批,却忽视输入可靠性、子群差异、阈值敏感性及部署可行性。本文提出RISED框架,通过BCa Bootstrap 95%置信区间、文献基准阈值和霍尔姆-邦弗朗尼校正,评估可靠性、包容性、敏感性、公平性与可部署性五个维度,其中公平性作为代理依赖诊断而非准入测试。应用于跨越35年七个队列(样本量303至99,492),发现传统AUROC无法揭示的问题:在Diabetes 130数据集中,可靠性虽达标(PSS=0.0004),但包容性(AUC差异=0.262)与敏感性(最大阈值翻转率49.1%)严重不达标;两个NHIS队列重现此现象。NHANES 2021–2023因特征完整获“不确定”结论;BRFSS 2024在移除高血压与胆固醇指标后,敏感性最差(最大阈值翻转率64.2%)。信用与收入预测队列也呈现相同模式,证实其跨领域普适性;多模型验证表明问题源于数据而非模型本身。RISED作为开源Python工具包,补充TRIPOD+AI、FUTURE-AI与Fairlearn标准所需的结构化证据。

原文摘要 · Abstract (English)

Clinical decision-support systems are expert systems whose recommendations clinicians act on directly, yet they are usually cleared on one aggregate accuracy number from a held-out test set. That number says nothing about input reliability under encoding shifts, subgroup gaps, threshold sensitivity, or operational feasibility. We present RISED, a pre-deployment evaluation framework operationalising five dimensions (Reliability, Inclusivity, Sensitivity, Equity, Deployability) through BCa bootstrap 95% confidence intervals, literature-grounded thresholds, and Holm-Bonferroni-corrected PASS / FAIL / INCONCLUSIVE verdicts; Equity is a proxy-dependence diagnostic rather than a gating test. Applied to seven cohorts spanning 35 years (n from 303 to 99,492), RISED surfaces failures invisible to AUROC: on Diabetes 130, Reliability passes by three orders of magnitude (PSS = 0.0004) while Inclusivity (AUC parity gap = 0.262) and Sensitivity (max threshold-flip rate 49.1%) fail decisively; both NHIS cohorts reproduce this. NHANES 2021-2023, with a complete feature profile, achieves INCONCLUSIVE verdicts; BRFSS 2024 produces the suite's most severe Sensitivity failure (max threshold-flip rate 64.2%) after instrument rotation removed hypertension and cholesterol. The pattern recurs on credit- and income-prediction cohorts, confirming domain-agnosticity; a multi-model check shows the failures are data-driven, not model-specific. RISED ships as an open-source Python package complementing TRIPOD+AI, FUTURE-AI, and Fairlearn with the structured numerical evidence those standards require but do not prescribe.

AI医疗评估框架公平性可部署性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。