EQPO提升医疗AI公平性,让不同人群诊断更均衡。
EQPO: Equitable Group Relative Policy Optimization for Clinical Reasoning
- 通过动态重加权样本,平衡不同人群学习进度
- 在7个诊断任务上降低43.9%的F1标准差
- 无需标签也能发现隐含群体,适合临床部署
医疗AI虽诊断性能优异,但在不同人口群体间表现不均,尤其对少数群体不利。尽管多模态推理基础模型推动了临床诊断发展,但基于强化学习的后训练常放大多数群体主导数据中的偏见。本文提出平等组相对策略优化(EQPO),一种分层强化学习方法,通过自适应重加权样本,依据子群占比、任务难度和数据源,促进异质临床人群的均衡学习。由于真实临床数据中常缺乏人口统计标注,EQPO还采用无监督聚类恢复隐含子群体。在覆盖X-ray、CT、dermoscopy、mammography、ultrasound五种模态的7个诊断基准上,相比原始GRPO,EQPO将QoQ-Med3-8B的F1标准差降低43.9%,最大跨群体F1差距减少42.7%;在MedGemma-4B上,相较去偏基线,预测公平性差距缩小27.2%,且F1提升12.5%——即使无任何人口标签。训练轨迹显示,EQPO持续提升公平性,而基线方法随训练退化;发现的隐含群体稳定且与掩码人口属性对齐。本文还发布了EquiMedGemma-4B与EquiQoQ-Med3-8B,具备公平性感知能力的临床视觉语言模型,在达到最先进准确率的同时,显著缩小人口差异。
原文摘要 · Abstract (English)
Medical AI systems demonstrated impressive diagnostic performance, yet they routinely show uneven accuracy across demographic groups, disadvantaging underrepresented populations. Although multimodal reasoning foundation models have pushed clinical diagnosis forward, reinforcement learning-based post-training tends to absorb and magnify the biases present in majority-dominated training corpora. We propose Equitable Group Relative Policy Optimization (EQPO), a hierarchical reinforcement learning method that encourages balanced learning across heterogeneous clinical populations by adaptively reweighting samples according to subgroup representation, task difficulty, and data source. As demographic annotations are frequently missing in real-world clinical data, EQPO additionally applies unsupervised clustering to recover latent subpopulations when they are unavailable. On 7 diagnostic benchmarks covering 5 modalities (X-ray, CT, dermoscopy, mammography, ultrasound), EQPO reduces F1 standard deviation by 43.9% and the maximum cross-group F1 gap by 42.7% on QoQ-Med3-8B over vanilla GRPO, and narrows predictive parity gaps by 27.2% on MedGemma-4B over bias-mitigated RL baselines while raising F1 by 12.5% even without any demographic labels. Examining the training trajectory shows that EQPO steadily improves fairness over the course of optimization, in contrast to baseline methods whose fairness degrades as training proceeds, and the discovered implicit groups remain stable and align with masked demographic attributes. We further release EquiMedGemma-4B and EquiQoQ-Med3-8B, equitability-aware clinical VLLMs that attain state-of-the-art accuracy with markedly smaller demographic gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。