arXiv:2510.12071cs.LG2025-10被引 2

发现训练中样本影响会随学习阶段动态变化,甚至反转方向。

Influence Dynamics and Stagewise Data Attribution

  • 基于奇异学习理论构建分阶段数据归因框架
  • 实证发现影响可非单调变化,含符号反转与突变峰值
  • 适用于理解大模型学习过程中的语义层次演化

现有训练数据归因(TDA)方法将样本间影响视为静态,但神经网络在不同学习阶段表现出动态的影响模式。本文提出一种基于奇异学习理论的分阶段数据归因框架,预测影响可能非单调变化,包括在发展转折点出现符号反转和尖峰。我们首先在简化模型中通过分析与实验验证了这些预测,发现影响的动态变化直接映射到模型对语义层级的逐步学习过程。最后,我们在语言模型中大规模验证了该现象,发现词元级影响变化与已知的学习发展阶段高度吻合。

原文摘要 · Abstract (English)

Current training data attribution (TDA) methods treat the influence one sample has on another as static, but neural networks learn in distinct stages that exhibit changing patterns of influence. In this work, we introduce a framework for stagewise data attribution grounded in singular learning theory. We predict that influence can change non-monotonically, including sign flips and sharp peaks at developmental transitions. We first validate these predictions analytically and empirically in a toy model, showing that dynamic shifts in influence directly map to the model's progressive learning of a semantic hierarchy. Finally, we demonstrate these phenomena at scale in language models, where token-level influence changes align with known developmental stages.

数据归因学习动态语言模型奇异学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。