arXiv:2505.08345cs.LGcs.AI2025-05中稿 · ACM FAccT 2025被引 16

数据工程方式会显著影响解释结果,可能隐藏歧视问题。

SHAP-based Explanations are Sensitive to Feature Representation

  • 用不同方式表示特征(如年龄用直方图)会改变重要性评分。
  • 常见数据处理方法可使SHAP等解释工具产生误导性结果。
  • 研究揭示了解释可靠性受数据表示影响,适合关注AI公平性的读者。

局部特征解释是XAI工具包的核心组成部分,其通过“可解释”的特征表示计算特征重要性。在表格数据中,特征值本身常被视为可解释。本文探讨了数据工程选择对局部特征解释的影响。我们证明,常见的简单数据工程方法,如将年龄表示为直方图或以特定方式编码种族,可操控主流方法(如SHAP)得出的特征重要性。值得注意的是,解释对特征表示的敏感性可被攻击者利用,以掩盖歧视等问题。尽管这一现象直观,但系统性研究仍不足。以往工作集中于通过偏置数据或操纵模型来攻击解释器,据我们所知,这是首项揭示标准、看似无害的数据工程技巧也能误导解释器的研究。

原文摘要 · Abstract (English)

Local feature-based explanations are a key component of the XAI toolkit. These explanations compute feature importance values relative to an ``interpretable'' feature representation. In tabular data, feature values themselves are often considered interpretable. This paper examines the impact of data engineering choices on local feature-based explanations. We demonstrate that simple, common data engineering techniques, such as representing age with a histogram or encoding race in a specific way, can manipulate feature importance as determined by popular methods like SHAP. Notably, the sensitivity of explanations to feature representation can be exploited by adversaries to obscure issues like discrimination. While the intuition behind these results is straightforward, their systematic exploration has been lacking. Previous work has focused on adversarial attacks on feature-based explainers by biasing data or manipulating models. To the best of our knowledge, this is the first study demonstrating that explainers can be misled by standard, seemingly innocuous data engineering techniques.

可解释性数据工程公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。