用数学理论解释大模型推理为何会失效,揭示位置编码与模型深度的关键影响。
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
- 通过最优传输理论量化领域偏移,用Wasserstein距离衡量分布变化。
- 证明位置编码若不平移不变,会导致风险恒定,而旋转位置编码可控制误差。
- 首次给出TC⁰ Transformer的电路深度下界,说明仅扩宽模型无法克服局限。
尽管大语言模型推理的实证缩放定律已被广泛记录,但其在分布外(OOD)泛化背后的理论机制仍不明确。本文通过最优传输理论形式化推理过程,将离散轨迹投影至连续度量空间,利用Wasserstein-1距离量化领域偏移。基于Kantorovich对偶性,我们通过架构的Lipschitz连续性和函数逼近极限界定了OOD泛化能力。研究揭示两大约束:其一,依赖位置的注意力机制(如绝对位置编码)破坏平移不变性,导致Ω(1)的Lipschitz常数和期望风险;而平移不变机制(如旋转编码)能保持等变性并限制误差。其二,通过将序列回溯映射为Dyck-k语言,建立了TC⁰ Transformer的严格电路深度下界。增加物理层深度是避免表征坍塌的必要条件——这一限制无法通过扩大表示宽度突破,因Barron空间中存在不可约的逼近边界。在54种Transformer配置的组合搜索任务上的评估验证了这些界限,显示泛化风险随Wasserstein域偏移单调上升。
原文摘要 · Abstract (English)
While empirical scaling laws for LLM reasoning are well-documented, the theoretical mechanisms governing out-of-distribution (OOD) generalization remain elusive. We formalize reasoning via optimal transport, projecting discrete trajectories into a continuous metric space to quantify domain shifts using the Wasserstein-1 distance. Invoking Kantorovich duality, we bound OOD generalization via architectural Lipschitz continuity and functional approximation limits. This exposes two primary constraints. First, position-dependent attention (e.g., Absolute Positional Encoding) fails to preserve shift invariance, yielding an $Ω(1)$ Lipschitz constant and expected risk, whereas shift-invariant mechanisms (e.g., Rotary Embeddings) preserve equivariance and bound the error. Second, by mapping sequential backtracking to a Dyck-$k$ language, we establish a strict circuit depth lower bound for $\text{TC}^0$ Transformers. Scaling physical layer depth is necessary to avert representation collapse -- a constraint that scaling representation width cannot bypass due to irreducible approximation bounds in Barron spaces. Evaluations across 54 Transformer configurations on combinatorial search corroborate these bounds, demonstrating that generalization risk degrades monotonically with the Wasserstein domain shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。