提出解释一致性评分,检测糖尿病视网膜病变模型在不同人群间是否用相同视觉依据决策。
Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening

- 用JS散度量化不同群体的特征重要性图相似性,评估解释一致性
- 跨种族预测性能有差异,但解释一致性仍较高且与性能无关
- 提醒需超越准确率,关注模型决策依据是否一致,适合医疗AI公平性研究者
医疗影像公平性常通过子群表现指标评估,但模型在不同人口群体中是否依赖一致的视觉证据尚不明确。本文提出解释一致性评分(ECS),基于杰恩森-香农散度,量化不同子群间归因图的相似性。以糖尿病视网膜病变筛查为案例,全局及按疾病严重程度评估了ECS。实验显示,尽管不同种族间预测性能存在差异,解释一致性仍相对较高,且与性能差距无显著关联。结果表明,预测公平性与解释一致性反映模型行为的不同维度,推动公平性评估从仅关注预测性能扩展至更深层决策机制。
原文摘要 · Abstract (English)
Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consistent visual evidence across demographic groups. This work introduces the Explanation Consistency Score (ECS), a fairness-aware metric based on Jensen-Shannon divergence that quantifies the similarity of attribution maps across subgroups. Using diabetic retinopathy screening as a case study, ECS is evaluated globally and within disease severity. Experiments reveal that while predictive performance differs across ethnic groups, explanation consistency remains relatively high and shows no significant association with performance disparities. These findings suggest that predictive fairness and explanation consistency capture complementary dimensions of model behavior, motivating fairness evaluations that extend beyond predictive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。