用少量数据生成高质量公平医疗数据,降低门槛并提升隐私保护
FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples
- 基于大模型上下文学习与提示工程生成表格数据
- 仅用1%数据实现50%公平性提升,预测性能接近真实数据
- 适合医疗数据隐私受限场景下的研究者使用
合成医疗数据生成为解决临床研究中因隐私和监管限制带来的困境提供了潜在方案。然而,现有方法需掌握生成模型训练知识且计算资源消耗高。本文提出FairTabGen,一种基于大语言模型的表格数据生成框架,仅需原始数据的小样本即可生成高质量合成医疗数据。该方法结合上下文学习、提示优化与嵌入结构约束进行数据合成。在MIMIC-IV数据集上评估显示,本方法仅使用99%更少的数据量,公平性(无意识)提升50%,同时保持优异的预测效用。但发现种族群体数据分布存在偏差,影响人口均等性;随后在预处理阶段引入偏见缓解算法,使整体公平性进一步提升10%,验证了方法的有效性。
原文摘要 · Abstract (English)
Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require specialized knowledge about training generative models and require high computational resources. In this paper, we propose FairTabGen, an LLM-based tabular data generation framework that produces high-quality synthetic healthcare data using only a small subset of the original dataset. Our method combines in-context learning, prompt curation and embedding structural constraints for data synthesis. We evaluate performance on MIMIC-IV dataset. Our method using 99% less data and achieving 50% improvement for fairness through unawareness while maintaining competitive predictive utility. However, we observe data distribution of racial groups is skewed affecting demographic parity. We thereafter apply bias mitigation algorithms in the pre-processing stage, improving overall fairness by 10% highlighting effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。