arXiv:2606.00563cs.LGcs.AI2026-06KDD

提出可实用的偏差上限,帮助医生评估医疗模型在真实人群中的表现风险。

A Practical Upper Bound on Selection Bias Effects in Medical Prediction Models

论文配图:A Practical Upper Bound on Selection Bias Effects in Medical Prediction Models
图 1 · 摘自论文原文
  • 基于部分可观测数据构建偏差影响的理论上限
  • 在合成、半合成及MIMIC-IV真实数据上验证有效性
  • 适合医疗AI从业者部署前评估模型可靠性

选择偏差是真实世界数据中常见且难以避免的问题,会影响机器学习模型的泛化能力。当在有偏数据上训练的模型部署到更广泛的目标人群时,可能因泛化性能差而造成实际危害,尤其在医疗等高风险场景中尤为严重。这凸显了从业者在部署前可靠评估模型泛化能力的必要性。然而,现有预测模型性能的方法通常依赖于对目标分布的完全访问或对偏差生成机制的精确了解,这些假设在现实中不成立。为此,本文提出一种在现实条件下(仅部分观测选择机制与目标人群数据)下,针对目标人群中最坏情况模型性能的新型上界估计方法。通过在全合成数据、基于All of Us研究计划的半合成数据以及MIMIC-IV的真实选择偏差数据上的实验,验证了该方法的有效性与实用性。本工作为在原本不可行的设定下估计选择偏差影响提供了原则性且实用的工具,有助于医疗等领域从业者构建更安全、更具泛化能力的模型。

原文摘要 · Abstract (English)

Selection bias is a common and often unavoidable aspect of real-world data that challenges the generalizability of machine learning models. When models trained on biased data are deployed in the broader target population, poor model generalization may lead to real harm, particularly in high-risk settings such as healthcare. This risk highlights the need for practitioners to reliably assess model generalizability prior to deployment. However, existing methods for predicting model performance rely on unrealistic access to the target distribution or knowledge of the selection mechanism causing bias. To address these limitations, we propose a novel upper bound on the worst-case model performance on the target population under the realistic setting where the selection mechanism and the target population data are only partially observed. We demonstrate the validity and practical utility of our method through experiments on fully synthetic data, semi-synthetic data derived from the All of Us Research Program, and real-world selection bias in MIMIC-IV. Our work offers a principled and practical tool to estimate the impact of selection bias in an otherwise intractable setting, thereby enabling practitioners to build safer and more generalizable models in healthcare and beyond.

医疗AI偏差分析模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。