检验影响函数在大模型中的有效性,发现其普遍表现不佳。
Do Influence Functions Work on Large Language Models?
- 通过系统实验评估影响函数在多任务下的表现。
- 多数场景下影响函数预测准确率低,误差源于参数近似与收敛不确定性。
- 适合关注大模型可解释性及样本重要性分析的研究者参考。
影响函数用于量化训练数据点对模型预测的影响。尽管在传统机器学习中研究广泛,但在大语言模型(LLMs)中的应用仍有限。本文系统研究了影响函数在大模型中的有效性,发现其在多数任务中表现不佳。深入分析表明,性能差的原因包括:(1) 由于大模型规模导致估计iHVP分量时不可避免的近似误差;(2) 微调过程中的收敛不确定性;(3) 影响函数定义本身的问题——模型参数变化不必然对应行为变化。因此,本研究呼吁探索替代方法以识别关键训练样本。
原文摘要 · Abstract (English)
Influence functions are important for quantifying the impact of individual training data points on a model's predictions. Although extensive research has been conducted on influence functions in traditional machine learning models, their application to large language models (LLMs) has been limited. In this work, we conduct a systematic study to address a key question: do influence functions work on LLMs? Specifically, we evaluate influence functions across multiple tasks and find that they consistently perform poorly in most settings. Our further investigation reveals that their poor performance can be attributed to: (1) inevitable approximation errors when estimating the iHVP component due to the scale of LLMs, (2) uncertain convergence during fine-tuning, and, more fundamentally, (3) the definition itself, as changes in model parameters do not necessarily correlate with changes in LLM behavior. Thus, our study suggests the need for alternative approaches for identifying influential samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。