动态数据中解释会失效,新方法可低成本修复旧解释。
Counterfactual Explanations Under Concept Drift
- 用局部采样估算解释有效性与合理性方向,轻量更新已有反事实解释。
- 实验显示旧解释在数据漂移下迅速失效,维护后仍保持有效性和合理性。
- 适合需要持续可信解释的在线学习系统,如金融风控、医疗决策。
反事实解释(CFEs)提供可操作的补救路径,但现有方法假设数据与模型静态不变。在数据流等持续演化的环境中,模型随概念漂移不断更新,导致生成时有效的解释可能悄然失效,包括鲁棒型解释也未考虑连续漂移。本文提出一种轻量级、模型无关的更新方案:通过局部采样估计解释的有效性与合理性方向,在保持与原样本接近的前提下修复旧解释。在合成漂移数据流上的实验表明,初始生成的解释迅速失效,而维护后的解释能长期保持有效性与局部合理性,成本远低于重新生成。
原文摘要 · Abstract (English)
Counterfactual explanations (CFEs) provide actionable recourse, but most methods assume a static framework with fixed data and a trained classifier. This assumption breaks in evolving data environments, such as data streams, where online models are repeatedly updated under concept drift. We identify CFE maintenance in this setting as a previously overlooked problem: explanations that are valid when generated may silently become invalid as the model evolves, including robust CFEs, which are not designed for continuous drift. We propose a lightweight, model-agnostic update scheme that repairs existing CFEs using local sampling to estimate validity and plausibility directions while preserving proximity to the original instance. Experiments on synthetic drifting streams show that initially created CFEs rapidly lose validity, whereas maintained CFEs preserve validity and local plausibility at a lower cost than repeated regeneration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。