arXiv:2609.03533cs.LG2026-09

提出耦合缩放框架,解释模型如何因架构与优化差异而不同地利用数据结构。

Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws

  • 用任务结构与可访问几何的耦合关系解释有限预算下的缩放规律
  • 残差指数范围由架构-优化系统的覆盖速率与任务尾部衰减速率决定
  • 适用于分析模型缩放行为差异,适合研究者和系统设计者参考

现有理论基于数据几何或指定的数据-模型谱推导神经网络缩放规律,但相同数据上训练的系统在架构或优化变化时可能表现出不同的缩放特性,因其能有效到达的表示空间不同。本文提出耦合缩放框架,指出有限预算下的缩放依赖于任务结构与架构-优化系统所能达到的几何结构之间的关系。在可解的模式截断模型中,损失可分解为超出架构支持范围的目标能量和未解决的支持尾部。对于任意优先级顺序,残差介于最佳前N个支持尾部与最大完成高价值前缀之外的尾部之间。若累积尾部与覆盖率对数速率分别为γ_{A,T}和ρ_{A,O,T},则残差指数位于[ρ_{A,O,T}γ_{A,T}, γ_{A,T}]区间。在前缀外收益有界时,已完成前缀决定速率,α_{A,O,T}=ρ_{A,O,T}γ_{A,T};当a_{A,T,j}∼j^{-b_{A,T}}时,得α_{A,O,T}=ρ_{A,O,T}(b_{A,T}-1)。固定核特化从任务加权谱测度的零附近尾部推导训练时间指数,该测度独立于损失拟合。该框架将架构支持与有限预算获取相分离,并提出两项检验:静态任务相关几何应在相同预算下跟踪损失变化;多尺度几何应跟踪特定耦合的指数排序,包括跨不同任务的反转现象。对已发布涌现轨迹的审计识别出直接因子测试所需的控制条件,以分别测量几何与缩放拟合。

原文摘要 · Abstract (English)

Existing theories derive neural scaling from data geometry or a specified data-model spectrum, but systems trained on the same data can scale differently when architecture or optimization changes the representations they can efficiently reach. We introduce Coupled Scaling, a task-conditioned framework in which finite-budget scaling depends on the relation between task structure and the geometry accessible to an architecture-optimization system. In a solvable mode-truncation model, loss separates into target energy outside architectural support and an unresolved supported tail. For an arbitrary priority order, the residual lies between the best-N supported tail and the tail beyond the largest completed high-value prefix. If the cumulative-tail and coverage log-rates are $\gamma_{A,T}$ and $\rho_{A,O,T}$, the residual exponent lies in $[\rho_{A,O,T}\gamma_{A,T},\gamma_{A,T}]$. Under bounded off-prefix gain, the completed prefix is rate-determining and $\alpha_{A,O,T}=\rho_{A,O,T}\gamma_{A,T}$; for $a_{A,T,j}\asymp j^{-b_{A,T}}$, this gives $\alpha_{A,O,T}=\rho_{A,O,T}(b_{A,T}-1)$. A fixed-kernel specialization derives the training-time exponent from the near-zero tail of a task-weighted spectral measure defined independently of the loss fit. The framework separates architectural support from finite-budget acquisition and motivates two tests: static task-relevant geometry should track loss at a common budget, while multiscale geometry should track coupling-specific exponent ordering, including reversal across contrasting tasks. An audit of released emergence trajectories identifies the controls needed for a direct factorial test that measures geometry separately from the scaling fit.

缩放定律表示学习架构分析优化机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。