用增强方法提升生成医疗数据的公平性,避免特定人群被低估或高估。
MedEqualizer: A Framework Investigating Bias in Synthetic Medical Data and Mitigation via Augmentation
- 通过增强欠代表群体数据,改进生成模型的公平性。
- 在MIMIC-III数据上发现多个群体在合成数据中存在显著比例失衡。
- 框架不依赖具体模型,适合各类生成式医疗数据研究使用。
合成医疗数据生成为克服真实医疗数据的局限性、提升数据可及性提供了可行路径。然而,在受保护属性上的公平性对避免临床研究与决策中的偏见至关重要。本研究基于GAN模型,利用MIMIC-III数据集评估合成数据在人口统计学属性上的公平性,采用对数差异度量分析子群代表性。结果发现,许多子群在合成数据中存在显著过采或欠采现象。为此,我们提出MedEqualizer——一种模型无关的数据增强框架,在生成前对欠代表群体进行数据补充。实验表明,该方法显著提升了合成数据的人口结构平衡性,为实现更公平、更具代表性的医疗数据合成提供了有效方案。
原文摘要 · Abstract (English)
Synthetic healthcare data generation presents a viable approach to enhance data accessibility and support research by overcoming limitations associated with real-world medical datasets. However, ensuring fairness across protected attributes in synthetic data is critical to avoid biased or misleading results in clinical research and decision-making. In this study, we assess the fairness of synthetic data generated by multiple generative adversarial network (GAN)-based models using the MIMIC-III dataset, with a focus on representativeness across protected demographic attributes. We measure subgroup representation using the logarithmic disparity metric and observe significant imbalances, with many subgroups either underrepresented or overrepresented in the synthetic data, compared to the real data. To mitigate these disparities, we introduce MedEqualizer, a model-agnostic augmentation framework that enriches the underrepresented subgroups prior to synthetic data generation. Our results show that MedEqualizer significantly improves demographic balance in the resulting synthetic datasets, offering a viable path towards more equitable and representative healthcare data synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。