用错误率评估医保预授权系统公平性,避免因临床差异导致的误判。
Anterior's Approach to Fairness Evaluation of Automated Prior Authorization System
- 以模型错误率替代审批率作为公平性衡量标准,更符合医疗实际。
- 7166例病例分析显示,多数人群错误率一致,95%置信区间在±5%内。
- 适用于监管机构、医疗AI开发者,尤其关注算法合规性的人群。
医保预授权(PA)系统因人力不足和时效压力日益自动化。传统以审批率均等衡量公平性不适用,因临床指南与医疗必要性常随人口群体差异。本文提出基于模型错误率的公平性评估框架。基于7,166份人工评审病例及27项医疗必要性指南,评估了性别、年龄、种族/族裔、社会经济地位的差异。综合采用错误率对比、±5百分点容差带分析、统计功效评估及受控逻辑回归。多数群体错误率一致,置信区间在预设容差带内,无显著性能差异;种族/族裔群体点估计值小,但子组样本量有限,置信区间宽,检验功效不足,证据尚不明确。结果展示了一种严谨且契合监管要求的行政医疗AI公平性评估方法。
原文摘要 · Abstract (English)
Increasing staffing constraints and turnaround-time pressures in Prior authorization (PA) have led to increasing automation of decision systems to support PA review. Evaluating fairness in such systems poses unique challenges because legitimate clinical guidelines and medical necessity criteria often differ across demographic groups, making parity in approval rates an inappropriate fairness metric. We propose a fairness evaluation framework for prior authorization models based on model error rates rather than approval outcomes. Using 7,166 human-reviewed cases spanning 27 medical necessity guidelines, we assessed consistency in sex, age, race/ethnicity, and socioeconomic status. Our evaluation combined error-rate comparisons, tolerance-band analysis with a predefined $\pm$5 percentage-point margin, statistical power evaluation, and protocol-controlled logistic regression. Across most demographics, model error rates were consistent, and confidence intervals fell within the predefined tolerance band, indicating no meaningful performance differences. For race/ethnicity, point estimates remain small, but subgroup sample sizes were limited, resulting in wide confidence intervals and underpowered tests, with inconclusive evidence within the dataset we explored. These findings illustrate a rigorous and regulator-aligned approach to fairness evaluation in administrative healthcare AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。