arXiv:2412.01003cs.LGcs.CL2024-12被引 61

揭示大模型上下文学习中算法竞争动态,解释其行为的瞬变性。

Competition Dynamics Shape Algorithmic Phases of In-Context Learning

  • 构建马尔可夫链混合序列任务,统一研究上下文学习机制。
  • 发现四种算法在不同条件下交替主导模型行为,导致学习效果突变。
  • 适合关注大模型内在推理机制与实验条件敏感性的研究者。

上下文学习(ICL)显著拓展了大语言模型的通用能力,使其仅通过输入上下文即可适应新任务。已有研究多基于非序列化的简化设定,其结论普适性存疑。为此,我们提出一个合成的序列建模任务,模拟有限马尔可夫链混合过程。实验证明,该任务能复现多数已知的ICL现象,为研究提供统一框架。在此基础上,我们发现模型行为可分解为四类算法:基于模糊检索或推理,结合上下文的一元或二元统计。这些算法在竞争中动态主导行为,实验条件(如上下文大小、训练量)的变化会引发算法主导权的突变,揭示了ICL的瞬变本质。因此,我们主张将ICL视为多种算法的混合体,而非单一能力;跨所有设置的普遍性结论可能无法成立。

原文摘要 · Abstract (English)

In-Context Learning (ICL) has significantly expanded the general-purpose nature of large language models, allowing them to adapt to novel tasks using merely the inputted context. This has motivated a series of papers that analyze tractable synthetic domains and postulate precise mechanisms that may underlie ICL. However, the use of relatively distinct setups that often lack a sequence modeling nature to them makes it unclear how general the reported insights from such studies are. Motivated by this, we propose a synthetic sequence modeling task that involves learning to simulate a finite mixture of Markov chains. As we show, models trained on this task reproduce most well-known results on ICL, hence offering a unified setting for studying the concept. Building on this setup, we demonstrate we can explain a model's behavior by decomposing it into four broad algorithms that combine a fuzzy retrieval vs. inference approach with either unigram or bigram statistics of the context. These algorithms engage in a competition dynamics to dominate model behavior, with the precise experimental conditions dictating which algorithm ends up superseding others: e.g., we find merely varying context size or amount of training yields (at times sharp) transitions between which algorithm dictates the model behavior, revealing a mechanism that explains the transient nature of ICL. In this sense, we argue ICL is best thought of as a mixture of different algorithms, each with its own peculiarities, instead of a monolithic capability. This also implies that making general claims about ICL that hold universally across all settings may be infeasible.

上下文学习算法竞争大模型机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。