ML模型在糖尿病和心脏病预测中对老年人和女性表现更差
Disparate Model Performance and Stability in Machine Learning Clinical Support for Diabetes and Heart Diseases
- 用数据复杂度与模型任意性结合评估公平性
- 男性和年轻患者预测准确率更高,老年人结果不稳定
- 强调数据代表性不足会引发临床模型不公平
机器学习算法在生物医学信息学中对临床决策支持至关重要。然而,其预测性能在不同人口群体间存在差异,常因历史边缘化群体在训练数据中代表性不足所致。研究发现慢性病数据集及其衍生的机器学习模型普遍存在性别与年龄相关的不平等现象。为此提出一种新分析框架,结合系统性任意性与传统指标(如准确率)及数据复杂度。基于超过25,000名慢性病患者的分析显示,存在轻微的性别差异,男性预测准确率略高;显著的年龄差异表现为年轻患者准确率更高。值得注意的是,老年患者在七个数据集中表现出不一致的预测准确率,与更高的数据复杂度和更低的模型性能相关。这表明仅保证训练数据代表性并不足以实现公平结果,部署前必须解决模型任意性问题。
原文摘要 · Abstract (English)
Machine Learning (ML) algorithms are vital for supporting clinical decision-making in biomedical informatics. However, their predictive performance can vary across demographic groups, often due to the underrepresentation of historically marginalized populations in training datasets. The investigation reveals widespread sex- and age-related inequities in chronic disease datasets and their derived ML models. Thus, a novel analytical framework is introduced, combining systematic arbitrariness with traditional metrics like accuracy and data complexity. The analysis of data from over 25,000 individuals with chronic diseases revealed mild sex-related disparities, favoring predictive accuracy for males, and significant age-related differences, with better accuracy for younger patients. Notably, older patients showed inconsistent predictive accuracy across seven datasets, linked to higher data complexity and lower model performance. This highlights that representativeness in training data alone does not guarantee equitable outcomes, and model arbitrariness must be addressed before deploying models in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。