数据溯源方法对超参数敏感,新方法无需重训即可高效选参。
Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly Retraining
- 提出无需重训的轻量级超参数选择法,基于正则化项理论分析。
- 实验证明多数方法对关键超参数敏感,但传统验证成本过高。
- 适用于数据溯源、模型可解释性等需高效评估的场景。
数据溯源方法用于量化单个训练样本对模型的影响,在现代AI的数据中心应用中日益流行。尽管该领域新方法层出不穷,其超参数调优的影响仍缺乏深入研究。本文首次开展大规模实证研究,揭示多数数据溯源方法对特定关键超参数高度敏感。然而,与通常可通过廉价验证指标调参的机器学习算法不同,评估数据溯源性能往往需要在子集数据上重新训练模型,导致调参成本过高。为此,我们主张加强超参数行为的理论理解以指导高效调参策略。作为案例研究,我们对影响函数类方法中至关重要的正则化项进行理论分析,并据此提出一种无需模型重训的轻量级正则化值选择方法。该方法在多个标准数据溯源基准上验证有效。本研究识别出数据溯源实用化中的一个根本性但被忽视的挑战,强调未来方法开发中需重视超参数选择的讨论。
原文摘要 · Abstract (English)
Data attribution methods, which quantify the influence of individual training data points on a machine learning model, have gained increasing popularity in data-centric applications in modern AI. Despite a recent surge of new methods developed in this space, the impact of hyperparameter tuning in these methods remains under-explored. In this work, we present the first large-scale empirical study to understand the hyperparameter sensitivity of common data attribution methods. Our results show that most methods are indeed sensitive to certain key hyperparameters. However, unlike typical machine learning algorithms -- whose hyperparameters can be tuned using computationally-cheap validation metrics -- evaluating data attribution performance often requires retraining models on subsets of training data, making such metrics prohibitively costly for hyperparameter tuning. This poses a critical open challenge for the practical application of data attribution methods. To address this challenge, we advocate for better theoretical understandings of hyperparameter behavior to inform efficient tuning strategies. As a case study, we provide a theoretical analysis of the regularization term that is critical in many variants of influence function methods. Building on this analysis, we propose a lightweight procedure for selecting the regularization value without model retraining, and validate its effectiveness across a range of standard data attribution benchmarks. Overall, our study identifies a fundamental yet overlooked challenge in the practical application of data attribution, and highlights the importance of careful discussion on hyperparameter selection in future method development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。