公平性约束会显著改变模型解释,影响临床信任。
The Effect of Enforcing Fairness on Reshaping Explanations in Machine Learning Models
- 用Shapley值分析公平性改进对特征重要性的影响
- 在三个医疗数据集上发现特征排名变化明显且群体差异大
- 提示需同时评估准确率、公平性和可解释性
可信的医疗机器学习需兼具强预测性能、公平性与可解释性。尽管提升公平性可能影响预测表现已知,但其如何影响解释性——临床信任的关键——尚不明确。本文研究通过偏见缓解技术提升公平性后,对基于Shapley值的特征重要性排名的重塑效应。我们在三个数据集上评估:儿童尿路感染风险、直接抗凝药物出血风险及再犯风险。结果表明,跨种族子群体提升公平性会显著改变特征重要性排名,且变化模式存在群体差异。这强调了在模型评估中必须联合考虑准确性、公平性与可解释性,而非孤立看待。
原文摘要 · Abstract (English)
Trustworthy machine learning in healthcare requires strong predictive performance, fairness, and explanations. While it is known that improving fairness can affect predictive performance, little is known about how fairness improvements influence explainability, an essential ingredient for clinical trust. Clinicians may hesitate to rely on a model whose explanations shift after fairness constraints are applied. In this study, we examine how enhancing fairness through bias mitigation techniques reshapes Shapley-based feature rankings. We quantify changes in feature importance rankings after applying fairness constraints across three datasets: pediatric urinary tract infection risk, direct anticoagulant bleeding risk, and recidivism risk. We also evaluate multiple model classes on the stability of Shapley-based rankings. We find that increasing model fairness across racial subgroups can significantly alter feature importance rankings, sometimes in different ways across groups. These results highlight the need to jointly consider accuracy, fairness, and explainability in model assessment rather than in isolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。