同一架构模型因初始化不同,测试计算量差异由动态相变决定。
Dynamical phase selection controls compute scaling in looped transformers
- 通过迭代权重共享映射,将推理过程建模为动力学系统。
- 不同初始值导致相变类型不同,影响测试时计算量随难度的分布规律。
- 发现两种相:折叠相与尼马克-萨克型相,前者有明确计算缩放规律。
环状Transformer通过迭代权重共享映射执行推理,其计算成本由推理动力学决定。本文表明,尽管架构与目标相同、训练精度一致,但不同初始化会导致网络进入不同的动力学相,且相变临界点决定了测试时计算量的缩放方式。这些相由不同的分岔机制区分,包括鞍结折叠和尼马克-萨克型向有界非平稳运动的转变。在折叠相中,一维规范型还原可从训练映射的局部导数预测弛豫时间与谱隙幅度,得出无参数关系式 $τ(\varepsilon)[1-λ_{\max}(-\varepsilon)]\toπ$。结合问题难度的常规分布,相同的临界减速效应产生工作负载尾部 $P(τ>N)\sim N^{-2}$。在尼马克-萨克相中,折叠缩放律完全消失而非仅改变系数。因此,测试时计算量并非仅由架构决定,而是由训练所寻得解的动力学相主导。
原文摘要 · Abstract (English)
A looped transformer performs inference by iterating a weight-tied map, making its computation a dynamical process whose cost is set by the resulting inference dynamics. Here we show that networks with identical architecture and objective, trained to identical accuracy, nevertheless realize distinct dynamical phases depending strongly on initialization, and that the bifurcation defining each phase determines how test-time compute scales. The phases are distinguished by their bifurcation mechanisms, including a saddle-node fold and a Neimark-Sacker-type transition to bounded nonstationary motion. In the fold phase, a one-dimensional normal-form reduction predicts both the relaxation-time and spectral-gap amplitudes from local derivatives of the trained map, yielding the parameter-free relation $τ(\varepsilon)[1-λ_{\max}(-\varepsilon)]\toπ$. Composed with a regular distribution of problem difficulty, the same critical slowing down produces the workload-level tail $P(τ>N)\sim N^{-2}$. In the Neimark--Sacker phase, the fold scaling law disappears rather than merely changing its prefactor. Thus, test-time compute is not determined by architecture alone. It is governed by the dynamical phase of the solution found by training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。