arXiv:2505.20634cs.LGstat.ML2025-05

提出SGShift方法,精准定位导致模型性能下降的特征变化。

Explaining Concept Shift with Interpretable Feature Attribution

  • 将概念漂移建模为特征选择问题,识别源域与目标域间关键差异特征。
  • 在少量目标域样本下仍能准确识别漂移特征,优于基线方法。
  • 适用于医疗、时间序列等科学场景,帮助理解特征-标签关系演变。

概念漂移指特征条件下标签分布随领域变化,导致即使调优良好的机器学习模型在新领域也出现校准偏差。识别这些发生漂移的特征,有助于揭示不同领域间特征-标签关系的差异,尤其在时间、疾病状态、人群等科学维度上具有重要意义。本文提出SGShift方法,将概念漂移下的性能退化归因于一组稀疏的漂移特征。该方法将概念漂移视为特征选择任务,以学习源域与目标域模型性能差异的解释特征。这一框架可结合广义加性模型、敲诈法(knockoffs)和吸收法(absorption)等强大统计工具,有效识别漂移特征。我们在合成数据与真实数据上进行了广泛实验,涵盖多种机器学习模型,结果表明SGShift在少量目标域样本下即可更准确识别漂移特征,且对复杂概念漂移情形具有鲁棒性。

原文摘要 · Abstract (English)

Concept shift occurs when the distribution of labels conditioned on the features changes between domains, which can make even a well-tuned ML model miscalibrated on a new domain. Identifying these shifted features provides unique insight into how feature-label relationships differ between domains, considering the difference may be across a scientifically relevant dimension, such as time, disease status, population, etc. In this paper, we propose SGShift, a method for attributing performance degradation under concept shift in tabular data to a sparse set of shifted features. We frame concept shift as a feature selection task to learn the features that can explain performance differences between models in the source and target domain. This framework enables SGShift to adapt powerful statistical tools such as generalized additive models, knockoffs, and absorption towards identifying these shifted features. We conduct extensive experiments in synthetic and real data across various ML models and find SGShift can identify shifted features much more accurately than baseline methods, requires few samples in the shifted domain, and is robust to complex cases of concept shift.

概念漂移特征归因可解释性表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。