arXiv:2508.07297cs.LGcs.AI2025-08被引 2

用影响函数追溯模型预测来源,高效识别关键训练数据。

Revisiting Data Attribution for Influence Functions

  • 基于统计影响函数,无需重训快速估算数据点影响。
  • 可有效定位误导性样本,支持数据调试与模型问责。
  • 适合关注模型可解释性与数据质量的研究者。

数据溯源的目标是将模型的预测结果通过学习算法回溯至其训练数据,从而识别最具影响力的训练样本,并理解模型行为如何导致特定预测。理解单个训练样本对模型预测的影响,是机器学习可解释性、数据调试与模型问责的基础。影响函数源自稳健统计学,提供一种高效的首阶近似方法,可在不进行昂贵重训练的前提下,估计微调某个数据点权重或移除该点对模型参数及后续预测的影响。本文全面回顾了影响函数在深度学习中的数据溯源能力,探讨其理论基础、近期用于高效计算逆海森向量积的算法进展,并评估其在数据溯源与误标检测中的有效性。最后,指出现有挑战与未来在大规模真实场景中释放影响函数巨大潜力的可行方向。

原文摘要 · Abstract (English)

The goal of data attribution is to trace the model's predictions through the learning algorithm and back to its training data. thereby identifying the most influential training samples and understanding how the model's behavior leads to particular predictions. Understanding how individual training examples influence a model's predictions is fundamental for machine learning interpretability, data debugging, and model accountability. Influence functions, originating from robust statistics, offer an efficient, first-order approximation to estimate the impact of marginally upweighting or removing a data point on a model's learned parameters and its subsequent predictions, without the need for expensive retraining. This paper comprehensively reviews the data attribution capability of influence functions in deep learning. We discuss their theoretical foundations, recent algorithmic advances for efficient inverse-Hessian-vector product estimation, and evaluate their effectiveness for data attribution and mislabel detection. Finally, highlighting current challenges and promising directions for unleashing the huge potential of influence functions in large-scale, real-world deep learning scenarios.

可解释性影响函数数据溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。