模型通过内化思维链,从依赖提示到自主计算,提升学习效率。
Learning through Internalization

- 用Transformer内化思维链,将推理过程编码到权重中
- 在奇偶性任务上,内化后性能显著提升,且无需外部提示
- 揭示内化带来分布外性能下降的潜在风险,适合研究模型可解释性
我们研究神经网络系统如何通过内部化过程将显式计算步骤融入自身权重,并促进学习。重点考察Transformer如何内化半自动机模拟,特别是通过内化思维链(CoT)标记;分析哪些类别的半自动机更难内化,并揭示内化的副作用:分布外性能随内化程度逐步退化。我们首次提供了成功的内化可证明分析:在奇偶性学习任务中,简化的一层Transformer首先在显式思维链监督下学会目标,随后随着思维链标记逐步移除,内化为自回归生成方式,最终直接计算奇偶性。该任务在无思维链监督下难以从数据中学习。最后,讨论内化学习与近期提出的正分布偏移现象的关系。
原文摘要 · Abstract (English)
We study internalization processes, by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how they facilitate learning. We investigate how transformers internalize the simulation of semiautomata by internalizing chain-of-thought (CoT) tokens, which classes of semiautomata are harder to internalize, and expose the flip side of internalization, that is, a progressive degradation of out-of-distribution performance. We then provide the first provable analysis of successful internalization: for the task of learning parities, we show that a simplified one-layer transformer provably first learns the target with explicit CoT supervision and then internalizes the autoregressive generation as CoT tokens are progressively removed, learning to directly compute the parity. This task is computationally hard to learn from data without CoT supervision. Finally, we discuss how learning through internalization relates to the \textit{Positive Distribution Shift} phenomenon recently introduced by~\citet{Med+26}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。