arXiv:2509.09616cs.LGcs.AI2025-09被引 2

用群体反事实解释追踪模型决策变化,揭示概念漂移根源

Explaining Concept Drift through the Evolution of Group Counterfactuals

  • 通过分析群体反事实解释的演化轨迹,追踪决策逻辑变化
  • 能区分数据空间偏移与概念重定义等不同漂移成因
  • 适合需要可解释性诊断的动态系统开发者

动态环境中的机器学习模型常受概念漂移影响,数据分布变化导致性能下降。尽管漂移检测已较为成熟,但解释模型决策逻辑如何演变仍具挑战。本文提出一种新方法,通过分析基于群体的反事实解释(Group Counterfactual Explanations, GCEs)随时间的演变来解释概念漂移。该方法追踪漂移前后GCE聚类中心及其关联的反事实动作向量的变化,这些演化特征作为可解释代理,揭示模型决策边界及内在推理机制的结构性变化。我们构建了一个三层框架,融合数据层(分布偏移)、模型层(预测分歧)和提出的解释层,实现对漂移的全面诊断,可有效区分空间数据偏移与概念重标签等不同根本原因。

原文摘要 · Abstract (English)

Machine learning models in dynamic environments often suffer from concept drift, where changes in the data distribution degrade performance. While detecting this drift is a well-studied topic, explaining how and why the model's decision-making logic changes still remains a significant challenge. In this paper, we introduce a novel methodology to explain concept drift by analyzing the temporal evolution of group-based counterfactual explanations (GCEs). Our approach tracks shifts in the GCEs' cluster centroids and their associated counterfactual action vectors before and after a drift. These evolving GCEs act as an interpretable proxy, revealing structural changes in the model's decision boundary and its underlying rationale. We operationalize this analysis within a three-layer framework that synergistically combines insights from the data layer (distributional shifts), the model layer (prediction disagreement), and our proposed explanation layer. We show that such holistic view allows for a more comprehensive diagnosis of drift, making it possible to distinguish between different root causes, such as a spatial data shift versus a re-labeling of concepts.

概念漂移反事实解释可解释性动态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。