FairGen通过偏好对齐生成公平的医学图像,缓解数据偏差带来的诊断不公。
FairGen: Preference-Aligned Diffusion for Demographically Equitable Medical Image Synthesis

- 引入医生偏好指导扩散过程,提升不同人群图像覆盖率。
- 在皮肤、胸部和脑部影像上分别实现95.9%、80.0%、35.2%的公平性提升。
- 兼顾诊断准确率,适合医疗AI公平性研究与临床部署应用。
医学影像在现代诊疗中至关重要,人工智能正日益用于提升分析效率、准确性和医疗可及性。然而,医疗资源获取不均与疾病发病率差异导致临床影像数据存在严重的人口学失衡。此外,疾病在不同人口群体中表现特征各异,某些表型自然稀少。基于此类不平衡数据训练的AI模型可能加剧诊断偏见,扩大健康差距。本文提出FairGen,一种面向公平性的扩散框架,在保持病理性视觉特征的同时合成人口均衡的医学图像。通过嵌入医生对齐的偏好,该方法在生成阶段提升亚组覆盖度,并改善下游分类性能。在皮肤病学、放射学和神经影像基准任务中,FairGen分别实现皮肤图像95.9%、胸部X光80.0%、脑部MRI 35.2%的公平性提升,同时相对于原始临床数据训练的模型保持相当的诊断准确性。临床医生评审及独立队列外部验证表明,这些改进不仅超越标准保真度指标,且不限于原分布数据集。
原文摘要 · Abstract (English)
Medical imaging is central to modern diagnostics, and artificial intelligence (AI) systems are increasingly used to support image-based analysis by improving efficiency, accuracy, and access to care. However, inequities in healthcare access and differential disease prevalence create severe demographic imbalances in clinical image data. Such imbalances are compounded by the fact that diseases can manifest with distinct features across demographic groups, rendering certain phenotypic presentations naturally rare. AI models trained on such imbalanced data risk perpetuating diagnostic bias and widening healthcare disparities. Here we introduce FairGen, a fairness-aware diffusion framework that synthesizes demographically balanced medical images while preserving pathology-relevant visual features. By embedding physician-aligned preferences into the generation process, FairGen improves subgroup coverage during synthesis and downstream classification. Applied to dermatology, radiology, and neuroimaging benchmark tasks, FairGen achieves fairness improvements of 95.9% for skin images, 80.0% for chest radiography, and 35.2% for brain MRI, while maintaining competitive diagnostic accuracy relative to models trained on original clinical data. Clinician-facing expert review and external validation on independent cohorts further support that these gains extend beyond standard fidelity metrics and are not confined to the original in-distribution datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。