arXiv:2502.04375cs.CLcs.LG2025-02ICML被引 13

小初始化让大模型更爱推理,大初始化则偏向记忆。

An Analysis for Reasoning Bias of Language Models with Small Initialization

  • 通过调整参数初始值大小,调控模型偏好推理或记忆。
  • 小初始化使模型在推理任务上表现更好,提升明显。
  • 适合想优化模型推理能力的研究者参考。

基于Transformer的大语言模型在自然语言处理中表现出色。本研究探究参数初始化规模对模型训练行为和任务偏好的影响。发现较小的初始化规模促使模型倾向于推理任务,而较大的初始化规模则导致其偏好记忆类任务。通过真实数据集和精心设计的锚定函数验证了这一推理偏差。进一步分析初始训练动态表明,嵌入空间和自注意力机制等特定模块在塑造学习偏差中起关键作用。本文从模型训练动态角度提出理论框架加以解释,并在真实语言任务实验中证实了理论洞见。该工作深化了对初始化策略如何影响大模型推理性能的理解,为模型训练提供了重要指导。

原文摘要 · Abstract (English)

Transformer-based Large Language Models (LLMs) have revolutionized Natural Language Processing by demonstrating exceptional performance across diverse tasks. This study investigates the impact of the parameter initialization scale on the training behavior and task preferences of LLMs. We discover that smaller initialization scales encourage models to favor reasoning tasks, whereas larger initialization scales lead to a preference for memorization tasks. We validate this reasoning bias via real datasets and meticulously designed anchor functions. Further analysis of initial training dynamics suggests that specific model components, particularly the embedding space and self-attention mechanisms, play pivotal roles in shaping these learning biases. We provide a theoretical framework from the perspective of model training dynamics to explain these phenomena. Additionally, experiments on real-world language tasks corroborate our theoretical insights. This work enhances our understanding of how initialization strategies influence LLM performance on reasoning tasks and offers valuable guidelines for training models.

大模型推理偏差初始化训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。