arXiv:2601.22771stat.MLcs.LG2026-01

解决特征解释方法冲突问题,让不同解释结果更一致。

GRANITE: A Generalized Regional Framework for Identifying Agreement in Feature-Based Explanations

  • 将特征空间划分为交互影响小的区域,减少解释分歧。
  • 在真实数据集上验证,显著提升解释一致性。
  • 适合需要可信解释的AI应用开发者使用。

基于特征的解释方法旨在量化特征对模型行为的影响,但不同方法常产生矛盾结论。这种分歧主要源于特征交互处理方式和依赖关系建模差异。本文提出GRANITE,一种广义区域解释框架,通过将特征空间划分为交互与分布影响最小的区域,统一并优化现有方法。该框架整合了已有区域方法,扩展至特征组,并引入递归划分算法以估计这些区域。在真实数据集上的实验表明,GRANITE能有效提升解释的一致性与可解释性,为实现稳定可靠的特征解释提供实用工具。

原文摘要 · Abstract (English)

Feature-based explanation methods aim to quantify how features influence the model's behavior, either locally or globally, but different methods often disagree, producing conflicting explanations. This disagreement arises primarily from two sources: how feature interactions are handled and how feature dependencies are incorporated. We propose GRANITE, a generalized regional explanation framework that partitions the feature space into regions where interaction and distribution influences are minimized. This approach aligns different explanation methods, yielding more consistent and interpretable explanations. GRANITE unifies existing regional approaches, extends them to feature groups, and introduces a recursive partitioning algorithm to estimate such regions. We demonstrate its effectiveness on real-world datasets, providing a practical tool for consistent and interpretable feature explanations.

特征解释可解释AI一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。