首个融合大模型的因果公平数据生成方法,提升医疗数据公平性。
FairCauseSyn: Towards Causally Fair LLM-Augmented Synthetic Data Generation
- 用大模型增强生成,保留因果结构以提升公平性
- 生成数据在因果公平度上偏差小于10%
- 训练模型后敏感属性偏见降低70%,适合医疗研究
合成数据生成通过生成模型基于真实数据创建新数据。在医疗应用中,生成高质量且保持敏感属性公平性的数据对实现公平结果至关重要。现有基于GAN和大模型的方法主要关注反事实公平性,多用于金融与法律领域。因果公平性通过保留因果结构提供更全面的评估框架,但当前合成数据生成方法尚未在医疗场景中解决此问题。为此,我们提出首个结合大模型的合成数据生成方法,以增强因果公平性,使用真实世界表格型医疗数据。生成数据在因果公平性指标上偏离真实数据不足10%。使用因果公平预测器训练时,合成数据相较真实数据将敏感属性偏见降低70%。本工作提升了公平合成数据的可及性,支持更公平的医疗研究与医疗服务。
原文摘要 · Abstract (English)
Synthetic data generation creates data based on real-world data using generative models. In health applications, generating high-quality data while maintaining fairness for sensitive attributes is essential for equitable outcomes. Existing GAN-based and LLM-based methods focus on counterfactual fairness and are primarily applied in finance and legal domains. Causal fairness provides a more comprehensive evaluation framework by preserving causal structure, but current synthetic data generation methods do not address it in health settings. To fill this gap, we develop the first LLM-augmented synthetic data generation method to enhance causal fairness using real-world tabular health data. Our generated data deviates by less than 10% from real data on causal fairness metrics. When trained on causally fair predictors, synthetic data reduces bias on the sensitive attribute by 70% compared to real data. This work improves access to fair synthetic data, supporting equitable health research and healthcare delivery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。