用因果推理生成更真实、高效的模型解释,让决策原因一目了然。
A New Approach to Backtracking Counterfactual Explanations: A Unified Causal Framework for Efficient Model Interpretability
- 基于反事实回溯构建因果框架,提升解释合理性
- 相比传统方法,计算开销更低且解释更贴近实际
- 适合需要可操作解释的高风险决策场景
反事实解释通过寻找能使模型输出不同的替代输入,提供对模型决策的局部洞察。然而,传统方法常忽略因果关系,导致解释不现实。尽管新方法引入因果性,但计算成本高昂。为此,我们提出一种名为 BRACE 的高效方法,基于反事实回溯并融合因果推理,生成可操作的解释。我们分析现有方法的局限性,提出新框架及其特性,并揭示其与旧方法的关系,证明在特定情况下可统一后者。实验表明,该方法能提供对模型输出更深入的理解。
原文摘要 · Abstract (English)
Counterfactual explanations enhance interpretability by identifying alternative inputs that produce different outputs, offering localized insights into model decisions. However, traditional methods often neglect causal relationships, leading to unrealistic examples. While newer approaches integrate causality, they are computationally expensive. To address these challenges, we propose an efficient method called BRACE based on backtracking counterfactuals that incorporates causal reasoning to generate actionable explanations. We first examine the limitations of existing methods and then introduce our novel approach and its features. We also explore the relationship between our method and previous techniques, demonstrating that it generalizes them in specific scenarios. Finally, experiments show that our method provides deeper insights into model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。