揭示语言模型预测中支持样本的内在作用与影响机制
On Support Samples of Next Word Prediction
- 基于表示定理识别促进或抑制预测的支持样本
- 支持样本属性可提前预判,非支持样本助于防过拟合
- 深层网络中非支持样本对表征学习至关重要
语言模型在复杂决策中表现优异,但其决策依据仍不清晰。本文聚焦语言模型中的数据驱动可解释性,研究下一个词预测任务。利用表示定理,我们识别出两类支持样本:促进或抑制特定预测的样本。研究发现,是否为支持样本是内在属性,甚至可在训练前预测。尽管非支持样本在直接预测中影响较小,但对防止过拟合、塑造泛化能力与表征学习至关重要。值得注意的是,非支持样本的重要性随网络深度增加,在深层中显著影响中间表征的形成。这些发现揭示了数据与模型决策之间的相互作用,为理解语言模型行为与可解释性提供了新视角。
原文摘要 · Abstract (English)
Language models excel in various tasks by making complex decisions, yet understanding the rationale behind these decisions remains a challenge. This paper investigates \emph{data-centric interpretability} in language models, focusing on the next-word prediction task. Using representer theorem, we identify two types of \emph{support samples}-those that either promote or deter specific predictions. Our findings reveal that being a support sample is an intrinsic property, predictable even before training begins. Additionally, while non-support samples are less influential in direct predictions, they play a critical role in preventing overfitting and shaping generalization and representation learning. Notably, the importance of non-support samples increases in deeper layers, suggesting their significant role in intermediate representation formation. These insights shed light on the interplay between data and model decisions, offering a new dimension to understanding language model behavior and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。