arXiv:2512.00976cs.LGstat.OT2025-12

现有超声心动图数据集缺乏性别种族多样性,模型对少数群体预测能力存疑。

Subgroup Validity in Machine Learning for Echocardiogram Data

  • 分析六大数据集,发现性别、种族、民族信息缺失严重
  • 对两种主动脉瓣狭窄模型的子群分析显示验证证据不足
  • 呼吁提升少数群体数据量与标签报告质量,确保模型公平性

超声心动图数据集可用于训练深度学习模型以自动化解读心脏超声图像,从而扩大精准诊断影像的可及性。然而,这些数据集中患者的性别、种族和民族信息普遍报告不全,且未评估子群特定的预测性能。这种报告缺陷引发了子群有效性问题,必须在模型部署前解决。本文表明,当前公开的超声心动图数据集无法缓解子群有效性担忧。我们改进了两个数据集(TMED-2 和 MIMIC-IV-ECHO)的社会人口学报告。对六个公开数据集的分析显示,缺乏对性别多元患者的关注,且多数种族和民族群体样本量不足。进一步对两个已发表的主动脉瓣狭窄检测模型在 TMED-2 上进行探索性子群分析,发现对性别、种族和民族子群均缺乏足够的有效性证据。研究结果强调,未来工作需增加少数群体数据、改善人口统计信息报告,并开展以子群为中心的分析,以证明模型的子群有效性。

原文摘要 · Abstract (English)

Echocardiogram datasets enable training deep learning models to automate interpretation of cardiac ultrasound, thereby expanding access to accurate readings of diagnostically-useful images. However, the gender, sex, race, and ethnicity of the patients in these datasets are underreported and subgroup-specific predictive performance is unevaluated. These reporting deficiencies raise concerns about subgroup validity that must be studied and addressed before model deployment. In this paper, we show that current open echocardiogram datasets are unable to assuage subgroup validity concerns. We improve sociodemographic reporting for two datasets: TMED-2 and MIMIC-IV-ECHO. Analysis of six open datasets reveals no consideration of gender-diverse patients and insufficient patient counts for many racial and ethnic groups. We further perform an exploratory subgroup analysis of two published aortic stenosis detection models on TMED-2. We find insufficient evidence for subgroup validity for sex, racial, and ethnic subgroups. Our findings highlight that more data for underrepresented subgroups, improved demographic reporting, and subgroup-focused analyses are needed to prove subgroup validity in future work.

医疗AI子群公平数据偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。