arXiv:2503.05631cs.LG2025-03ICML被引 26

模型学习能力会随训练时间变化,出现消失又重现的现象。

Strategy Coopetition Explains the Emergence and Transience of In-Context Learning

  • 发现模型在长训练后转向一种混合策略(CIWL),取代了原有的上下文学习能力。
  • 上下文学习与新策略共用底层电路,形成既竞争又合作的动态关系。
  • 提出最小数学模型解释该现象,帮助设计让学习能力持久存在的训练方法。

上下文学习(ICL)是变换器模型中一种无需权重更新即可从上下文中学习的强大能力。近期研究发现,这种能力是一种短暂现象,可能在长时间训练后消失。本文旨在揭示其动态机制:我们发现,当ICL消失后,模型最终采用一种混合策略——‘上下文约束的权重学习’(CIWL),它与ICL竞争并最终主导模型行为,导致ICL的瞬时性。然而,两者共享部分子电路,表现出协同作用。例如,在我们的设定中,ICL无法独立快速出现,必须依赖于缓慢发展的渐近式CIWL的同步演化。因此,两者既竞争又合作,我们称之为‘策略共竞’(strategy coopetition)。我们提出了一个最小数学模型,复现了这些关键动态。基于该模型,我们识别出一种使ICL真正涌现且持久的设置。

原文摘要 · Abstract (English)

In-context learning (ICL) is a powerful ability that emerges in transformer models, enabling them to learn from context without weight updates. Recent work has established emergent ICL as a transient phenomenon that can sometimes disappear after long training times. In this work, we sought a mechanistic understanding of these transient dynamics. Firstly, we find that, after the disappearance of ICL, the asymptotic strategy is a remarkable hybrid between in-weights and in-context learning, which we term "context-constrained in-weights learning" (CIWL). CIWL is in competition with ICL, and eventually replaces it as the dominant strategy of the model (thus leading to ICL transience). However, we also find that the two competing strategies actually share sub-circuits, which gives rise to cooperative dynamics as well. For example, in our setup, ICL is unable to emerge quickly on its own, and can only be enabled through the simultaneous slow development of asymptotic CIWL. CIWL thus both cooperates and competes with ICL, a phenomenon we term "strategy coopetition." We propose a minimal mathematical model that reproduces these key dynamics and interactions. Informed by this model, we were able to identify a setup where ICL is truly emergent and persistent.

上下文学习模型动态共竞机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。