用GPT-4生成分组合成数据,提升医疗模型对少数群体的公平性。
Improving Equity in Health Modeling with GPT4-Turbo Generated Synthetic Data: A Comparative Study
- 针对不同人群分别生成合成数据,增强代表性不足群体的样本
- 在MIMIC-IV和弗雷明汉队列上,合成数据使模型跨群体性能更均衡
- 无需指定具体群体也能达到相似效果,适合快速部署
目标:医疗数据集中不同人口群体的代表程度不一,导致机器学习模型对某些群体表现更优。一种有前景的解决方案是生成合成数据以缓解非代表性数据带来的偏见。方法:基于大语言模型(LLM)合成数据的新进展,我们构建了一个分组独立生成合成数据的流程。研究使用MIMIC-IV与弗雷明汉队列(Offspring and OMNI-1 Cohorts)数据集,通过提示GPT4-Turbo生成特定群体的合成数据,并提供训练样例与数据背景。我们开展探索性分析评估生成数据质量,并在下游机器学习任务中测试其作为训练数据增强的效果,重点关注模型在不同群体间的性能表现。结果:总体上,使用GPT4-Turbo生成的数据增益优于传统基线模型;但在多数实验中,是否指定群体生成对性能影响不大。结论:我们提出一种无需额外训练即可利用大语言模型生成分组合成数据的方法,有助于改善医疗模型中的公平性,推动健康公平。未来需进一步研究大语言模型生成数据在非代表性医疗数据中的适用条件。
原文摘要 · Abstract (English)
Objective. Demographic groups are often represented at different rates in medical datasets. These differences can create bias in machine learning algorithms, with higher levels of performance for better-represented groups. One promising solution to this problem is to generate synthetic data to mitigate potential adverse effects of non-representative data sets. Methods. We build on recent advances in LLM-based synthetic data generation to create a pipeline where the synthetic data is generated separately for each demographic group. We conduct our study using MIMIC-IV and Framingham "Offspring and OMNI-1 Cohorts" datasets. We prompt GPT4-Turbo to create group-specific data, providing training examples and the dataset context. An exploratory analysis is conducted to ascertain the quality of the generated data. We then evaluate the utility of the synthetic data for augmentation of a training dataset in a downstream machine learning task, focusing specifically on model performance metrics across groups. Results. The performance of GPT4-Turbo augmentation is generally superior but not always. In the majority of experiments our method outperforms standard modeling baselines, however, prompting GPT-4-Turbo to produce data specific to a group provides little to no additional benefit over a prompt that does not specify the group. Conclusion. We developed a method for using LLMs out-of-the-box to synthesize group-specific data to address imbalances in demographic representation in medical datasets. As another "tool in the toolbox", this method can improve model fairness and thus health equity. More research is needed to understand the conditions under which LLM generated synthetic data is useful for non-representative medical data sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。