提出通用医疗数据偏见检测框架,识别训练与测试数据中的隐藏偏差。
Detecting Dataset Bias in Medical AI: A Generalized and Modality-Agnostic Auditing Framework
- 通过分析标签与患者属性、采集环境的关系,量化数据偏见风险。
- 在皮肤病变、电子病历、重症监护三类数据中发现传统方法忽略的细微偏差。
- 适用于多模态医疗数据,帮助开发更可靠的AI系统。
人工智能正成为循证医学的核心。尽管在医疗领域取得诸多成功,但部署中仍频繁出现重大缺陷与意外行为。主要原因是AI依赖关联学习,非代表性数据集会放大或掩盖潜在偏见。为此,我们提出G-AUDIT:一种模态无关的数据集审计框架,可生成关于训练/测试数据偏见来源的针对性假设。该方法分析任务标签与患者属性(如年龄、性别)及环境/采集特征(如临床机构、成像协议)之间的关系,量化数据属性引发捷径学习的风险,或在测试阶段隐藏虚假关联预测的可能性。我们在三大不同模态和任务的大规模医疗数据集上验证了该方法:图像中的皮肤病变分类、电子病历中的污名化语言识别、重症监护的表格数据死亡率预测。在每种场景中,G-AUDIT均成功识别出传统定性方法常忽略的微妙偏见,凸显其在揭示数据层面风险、支持可靠AI系统开发方面的实用价值。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) is now firmly at the center of evidence-based medicine. Despite many success stories that edge the path of AI's rise in healthcare, there are comparably many reports of significant shortcomings and unexpected behavior of AI in deployment. A major reason for these limitations is AI's reliance on association-based learning, where non-representative machine learning datasets can amplify latent bias during training and/or hide it during testing. To unlock new tools capable of foreseeing and preventing such AI bias issues, we present G-AUDIT. Generalized Attribute Utility and Detectability-Induced bias Testing (G-AUDIT) for datasets is a modality-agnostic dataset auditing framework that allows for generating targeted hypotheses about sources of bias in training or testing data. Our method examines the relationship between task-level annotations (commonly referred to as ``labels'') and data properties including patient attributes (e.g., age, sex) and environment/acquisition characteristics (e.g., clinical site, imaging protocols). G-AUDIT quantifies the extent to which the observed data attributes pose a risk for shortcut learning, or in the case of testing data, might hide predictions made based on spurious associations. We demonstrate the broad applicability of our method by analyzing large-scale medical datasets for three distinct modalities and machine learning tasks: skin lesion classification in images, stigmatizing language classification in Electronic Health Records (EHR), and mortality prediction for ICU tabular data. In each setting, G-AUDIT successfully identifies subtle biases commonly overlooked by traditional qualitative methods, underscoring its practical value in exposing dataset-level risks and supporting the downstream development of reliable AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。