对比三种适配方法在胸片模型中的公平性表现,发现性能提升不等于公平性改善。
Subgroup performance analysis of adaptation strategies for chest X-ray foundation models

- 用线性头、MLP和注意力池化三种方式适配冻结的Rad-DINO模型
- 注意力池化效果最好但对种族编码最强,公平性未随性能提升而改善
- 模型层间差异与公平性无固定关系,需针对任务具体评估
基础模型在下游医学影像任务中日益被适配,但适配策略对子群体公平性的影响尚不明确。本文研究三种参数高效适配技术——基于原始CLS token的线性头、MLP和多层图像块特征上的注意力池化模块——在冻结的Rad-DINO胸片编码器上的表现。在MIMIC-CXR数据集上,于种族、性别和成像视角子群体中评估八种病理分类性能,使用保留流行率且人口统计平衡的测试集,并分析各适配器对受保护属性(如种族)的编码强度。结果表明:注意力池化实现最强整体判别性能且对种族编码最强烈,但整体性能提升并不一致降低子群体差异。值得注意的是,早期网络层对种族编码最弱,却产生最大性能差距。进一步探索不同注意力池化层组合发现,所选层、属性编码强度与子群体公平性之间无稳定关联。结果表明,更丰富的表征可提升准确率,但公平性影响高度依赖任务,不可仅从编码强度或总体性能推断,必须直接评估。
原文摘要 · Abstract (English)
Foundation models are increasingly adapted for downstream medical imaging tasks, yet the influence of the chosen adaptation strategy on subgroup fairness remains poorly understood. We investigate how three parameter-efficient adaptation techniques, including linear heads on the raw CLS token, an MLP, and an attention-pooling module over multi-layer patch features, affect both pathology classification performance and subgroup disparities when applied to the frozen Rad-DINO chest X-ray encoder. Using MIMIC-CXR, we evaluate eight pathologies across race, sex, and imaging-view subgroups on a prevalence-preserving, demographically balanced test set, and additionally probe how strongly each adapter encodes protected attributes. We find that attention pooling achieves the strongest overall discriminative performance and encodes attributes, particularly race, most strongly, but that improved overall performance does not consistently reduce subgroup disparities. Notably, stronger attribute encoding did not correspond to larger disparities: early network layers encoded race most weakly yet produced the largest subgroup performance gaps. Exploring different attention-pooling layer combinations further revealed no consistent relationship between the layers pooled, attribute encoding strength, and subgroup fairness. Our results indicate that richer, more expressive representations can improve accuracy while leaving fairness implications task-dependent and unpredictable, which must be assessed directly and per-task rather than inferred from encoding strength or overall performance alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。