arXiv:2602.01312cs.LG2026-02

TRAK算法虽有误差,但能准确保留数据重要性排序。

Imperfect Influence, Preserved Rankings: A Theory of TRAK for Data Attribution

  • 用核机器近似模型,结合留一法风险估算数据影响
  • 误差虽大,但数据相对重要性排序仍高度一致
  • 适合关注数据贡献排名的可解释性研究者

数据归因旨在将模型预测追溯至具体训练数据,是理解复杂AI模型的重要工具。广泛使用的TRAK算法首先用核机器近似模型,再利用留一法(ALO)风险的近似技术解决该问题。尽管其在实践中表现良好,但其理论基础及失效场景仍不明确。本文对TRAK算法进行理论分析,刻画其性能并量化其依赖近似带来的误差。结果表明,尽管近似引入显著误差,TRAK估计的影响值与原始影响仍高度相关,因而基本保持了数据点间的相对排序。通过大量模拟与实证研究验证了理论结论。

原文摘要 · Abstract (English)

Data attribution, tracing a model's prediction back to specific training data, is an important tool for interpreting sophisticated AI models. The widely used TRAK algorithm addresses this challenge by first approximating the underlying model with a kernel machine and then leveraging techniques developed for approximating the leave-one-out (ALO) risk. Despite its strong empirical performance, the theoretical conditions under which the TRAK approximations are accurate as well as the regimes in which they break down remain largely unexplored. In this paper, we provide a theoretical analysis of the TRAK algorithm, characterizing its performance and quantifying the errors introduced by the approximations on which the method relies. We show that although the approximations incur significant errors, TRAK's estimated influence remains highly correlated with the original influence and therefore largely preserves the relative ranking of data points. We corroborate our theoretical results through extensive simulations and empirical studies.

数据归因可解释性理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。