arXiv:2604.03478cs.LG2026-04

数据合并未必提升公平性,需结合模型校准才能有效改善医疗模型对子群体的表现。

Investigating Data Interventions for Subgroup Fairness: An ICU Case Study

  • 通过对比不同医院EHR数据合并策略,分析其对子群体性能的影响。
  • 数据量增加可能因分布偏移导致公平性下降,好数据不等于好结果。
  • 结合数据筛选与模型后处理,才能稳定提升子群体表现,适合医疗AI研究者。

在高风险医疗决策中,机器学习模型的算法偏见可能加剧对特定人群的系统性伤害,而这些偏见常源于训练数据。实际应用中,数据干预受限于可用数据源的质量。本文以eICU Collaborative Research Database和MIMIC-IV两个临床数据集为例,发现数据合并既可能提升也可能损害模型公平性与性能,且许多直观的数据选择策略不可靠。我们比较了基于模型的后处理校准与数据中心的添加策略,发现二者结合对改善子群体表现至关重要。研究质疑了‘更好数据即更公平’的传统观念,强调需协同使用数据与模型方法应对公平性挑战。

原文摘要 · Abstract (English)

In high-stakes settings where machine learning models are used to automate decision-making about individuals, the presence of algorithmic bias can exacerbate systemic harm to certain subgroups of people. These biases often stem from the underlying training data. In practice, interventions to "fix the data" depend on the actual additional data sources available -- where many are less than ideal. In these cases, the effects of data scaling on subgroup performance become volatile, as the improvements from increased sample size are counteracted by the introduction of distribution shifts in the training set. In this paper, we investigate the limitations of combining data sources to improve subgroup performance within the context of healthcare. Clinical models are commonly trained on datasets comprised of patient electronic health record (EHR) data from different hospitals or admission departments. Across two such datasets, the eICU Collaborative Research Database and the MIMIC-IV dataset, we find that data addition can both help and hurt model fairness and performance, and many intuitive strategies for data selection are unreliable. We compare model-based post-hoc calibration and data-centric addition strategies to find that the combination of both is important to improve subgroup performance. Our work questions the traditional dogma of "better data" for overcoming fairness challenges by comparing and combining data- and model-based approaches.

医疗AI公平性数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。