解决线性样条模型中边际Shapley值的外推问题,提升可解释性可靠性。
On Model Extrapolation in Marginal Shapley Values
- 采用边际平均法避免模型外推,结合因果信息还原真实贡献
- 在简单线性样条模型中验证外推会显著扭曲特征重要性估计
- 适用于需要高可靠性的实际数据场景,尤其关注因果解释的用户
随着复杂机器学习模型应用日益广泛,可靠的可解释性方法需求迫切。当前最流行的解释方法之一是基于Shapley值。计算Shapley值主要有两种常用方法:条件型与边际型,当特征相关时结果不同。此前研究指出,条件型方法因隐含因果假设而存在根本缺陷。然而,边际型方法常导致模型外推,使预测在未见区域不可靠。本文以简单线性样条模型为例,系统分析模型外推对边际Shapley值的影响,并提出一种新方法:在使用边际平均的同时避免外推,并通过引入因果信息复现因果型Shapley值。最后在真实数据上验证了该方法的有效性。
原文摘要 · Abstract (English)
As the use of complex machine learning models continues to grow, so does the need for reliable explainability methods. One of the most popular methods for model explainability is based on Shapley values. There are two most commonly used approaches to calculating Shapley values which produce different results when features are correlated, conditional and marginal. In our previous work, it was demonstrated that the conditional approach is fundamentally flawed due to implicit assumptions of causality. However, it is a well-known fact that marginal approach to calculating Shapley values leads to model extrapolation where it might not be well defined. In this paper we explore the impacts of model extrapolation on Shapley values in the case of a simple linear spline model. Furthermore, we propose an approach which while using marginal averaging avoids model extrapolation and with addition of causal information replicates causal Shapley values. Finally, we demonstrate our method on the real data example.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。