arXiv:2609.06173cs.LG2026-09

提出FMMO框架,揭示局部解释稳定但全局漂移的隐藏风险

FMMO: Detecting the Divergence Between Local Attribution and Global Drift

  • 结合全局代理模型与使用率监测,检测局部解释与整体漂移的背离
  • 发现保护群体误报率飙升时,局部归因仍保持稳定,传统XAI无法预警
  • 适合关注模型公平性监控的工程师和算法审计人员

部署后性能漂移对算法问责构成重大威胁,尤其在真实标签延迟时,性能下降可能成为“无声失败”。尽管可解释AI(XAI)常被用于审计此类变化,我们发现主流局部归因方法(如TreeSHAP)在模型可靠性崩溃时仍表现出误导性的稳定性。本文提出面向模型监控与可观测性的框架FMMO,旨在揭示局部解释稳定性与全局分布漂移之间的差异。在基准、合成及真实数据集上,我们验证了局部XAI方法无法识别漂移引发的不公平影响,特别是当保护群体的假阳性率上升而特征归因保持不变时。通过整合全局代理模型与模型使用度量,FMMO有效弥补了这一公平性盲区,使利益相关方能够察觉标准局部XAI工具所忽略的歧视性退化。

原文摘要 · Abstract (English)

Post-deployment drift poses a critical risk to algorithmic accountability, particularly when ground truth labels are delayed and performance degradation becomes a "silent failure". While Explainable AI (XAI) is often relied upon to audit these shifts, we demonstrate that popular local attribution methods (e.g., TreeSHAP) can exhibit misleading stability even as model reliability collapses. In this paper, we propose a Framework for Model Monitoring and Observability (FMMO) designed to expose the divergence between local explanation stability and global distribution shifts. Using benchmark, synthetic, and real-world datasets, we show that local XAI methods fail to flag drift-induced disparate impact, specifically where False Positive Rates spike for protected groups while feature attributions remain unchanged. By integrating global surrogate models with model utilization measurements, FMMO mitigates this fairness blind spot, ensuring that stakeholders can detect discriminatory deterioration that standard local XAI tools overlook.

模型监控可解释性公平性漂移检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。