arXiv:2507.11855cs.LG2025-07NeurIPS

区分特征值与位置对模型预测的影响,提升序列模型可解释性。

OrdShap: Feature Position Importance for Sequential Black-Box Models

  • 通过打乱特征顺序分析模型响应,分离位置与值的影响
  • 在医疗、自然语言等数据上验证了位置重要性可被精准捕捉
  • 理论严谨,适合关注序列模型决策逻辑的研究者

序列深度学习模型在具有时间或序列依赖性的领域表现优异,但其复杂性要求事后特征归因方法以理解预测过程。现有技术通常假设特征顺序固定,混淆了(1)特征值和(2)其在输入序列中的位置的双重影响。为此,我们提出OrdShap,一种新型归因方法,通过分析特征位置置换对模型预测的影响,实现两者的解耦。我们建立了OrdShap与Sanchez-Bergantiños值之间的博弈论关联,为位置敏感归因提供了理论基础。在医疗、自然语言处理及合成数据集上的实证结果表明,OrdShap能有效捕捉特征值与位置的重要性,并揭示模型行为的深层机制。

原文摘要 · Abstract (English)

Sequential deep learning models excel in domains with temporal or sequential dependencies, but their complexity necessitates post-hoc feature attribution methods for understanding their predictions. While existing techniques quantify feature importance, they inherently assume fixed feature ordering - conflating the effects of (1) feature values and (2) their positions within input sequences. To address this gap, we introduce OrdShap, a novel attribution method that disentangles these effects by quantifying how a model's predictions change in response to permuting feature position. We establish a game-theoretic connection between OrdShap and Sanchez-Bergantiños values, providing a theoretically grounded approach to position-sensitive attribution. Empirical results from health, natural language, and synthetic datasets highlight OrdShap's effectiveness in capturing feature value and feature position attributions, and provide deeper insight into model behavior.

可解释性序列模型特征归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。