用特征发现隐藏分组,比传统方法更早发现医疗模型性能差异。
Subgroup Performance Analysis in Hidden Stratifications
- 基于学习到的特征表示自动发现潜在患者分组。
- 新方法在胸部X光和皮肤病变分类中揭示了更大性能差异。
- 适合关注医疗AI可信度与公平性的研究者使用。
机器学习模型在不同患者群体间可能存在显著性能差异。仅通过元数据(如患者性别)进行传统子群分析,往往无法覆盖性能变化的主要原因。基于学习特征表示的子群发现技术有望揭示隐藏的分层结构,提供更细粒度的性能报告。然而,真实数据中缺乏真实分组标签,使得子群发现难以评估。本文首次将子群发现应用于胸部X光和皮肤病变分类中的性能监控,提出新颖评估策略,证明即使不依赖分类标签或元数据,简化版子群发现方法仍能暴露比传统方法更大的性能差异。这是首个有力证据表明,子群发现可成为医疗可信AI全面验证与监控的关键工具。
原文摘要 · Abstract (English)
Machine learning (ML) models may suffer from significant performance disparities between patient groups. Identifying such disparities by monitoring performance at a granular level is crucial for safely deploying ML to each patient. Traditional subgroup analysis based on metadata can expose performance disparities only if the available metadata (e.g., patient sex) sufficiently reflects the main reasons for performance variability, which is not common. Subgroup discovery techniques that identify cohesive subgroups based on learned feature representations appear as a potential solution: They could expose hidden stratifications and provide more granular subgroup performance reports. However, subgroup discovery is challenging to evaluate even as a standalone task, as ground truth stratification labels do not exist in real data. Subgroup discovery has thus neither been applied nor evaluated for the application of subgroup performance monitoring. Here, we apply subgroup discovery for performance monitoring in chest x-ray and skin lesion classification. We propose novel evaluation strategies and show that a simplified subgroup discovery method without access to classification labels or metadata can expose larger performance disparities than traditional metadata-based subgroup analysis. We provide the first compelling evidence that subgroup discovery can serve as an important tool for comprehensive performance validation and monitoring of trustworthy AI in medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。