arXiv:2608.25981cs.CVcs.AI2026-08

提出新框架FRAME,区分医学影像公平性中的采样差异与真实偏差。

FRAME: separating sampling variation from representational cause in medical imaging fairness

论文配图:FRAME: separating sampling variation from representational cause in medical imaging fairness
图 1 · 摘自论文原文
  • 构建公平参考基准,量化在当前样本量下纯随机导致的性能差异
  • 实测70万张图像中,41%的种族差异由采样误差解释,22%的年龄差异同理
  • 可帮助研究者判断哪些差异需机制分析,哪些仅因数据量不足产生

子组性能差异是医疗影像公平性偏见的标准证据,常规做法是移除模型对人口统计信息的编码。本文提出公平模型参考与机制评估框架(FRAME),分两步审计该主张:第一步生成公平参考,即在精确公平条件下,基于实际子组规模的差异分布;第二步在表示空间中使用两个算子测试剩余差异。一个算子从构造上无法改变组内排序,实验覆盖702,206张图像和36个编码器,参考基准解释了中位数41%的种族差异和22%的年龄差异。注入人口统计可解码性不改变剩余差异,而将群体与疾病方向纠缠使种族差异从0.077升至0.118。所有干预措施对剩余差异的影响均不超过随机种子变化。这些干预降低操作点差异,但组内排序差异中位数仍为0.000。应用于9篇已发表研究中的89个差异,参考基准解释了中位数25%的率差异和70%的ROC曲线下面积差异。图像-文本预训练则将最差组性能提升约0.05。在选择干预前应用FRAME,可区分需机制解释的差异与仅由采样变异引起的差异。

原文摘要 · Abstract (English)

Subgroup performance differences are the standard evidence for fairness bias in medical imaging, and the usual response removes the demographic information that a model encodes. Here we introduce Fair-model Reference And Mechanism Evaluation (FRAME), a two-step framework for auditing such a claim. The first step derives a fair-model reference, the distribution of the difference under exact fairness at the observed subgroup sizes. In the second step, we test the remainder with two operators in representation space. One operator cannot change a within-group ranking by construction. Across 702,206 images and 36 encoders, the reference accounts for a median 41% of the reported race difference and 22% of the age difference. Injecting demographic decodability leaves the remainder unchanged, while entangling the group with the disease direction raises the race difference from 0.077 to 0.118. No intervention we tested changes the remainder more than a change of random seed does. Those interventions reduce a difference at the operating point and leave the within-group ranking difference at a median of 0.000. Applied to 89 differences in 9 published studies across 6 medical imaging modalities, the reference accounts for a median 25% of a rate difference and 70% of a difference in the area under the receiver operating characteristic curve. Image-text pretraining instead raises worst-group performance by about 0.05. Applying FRAME before choosing an intervention could distinguish differences that need a mechanistic explanation from differences compatible with sampling variation at the current cohort sizes.

医疗影像公平性采样偏差机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。