arXiv:2607.14144cs.AIcs.IT2026-07被引 1

模型能力收敛依赖访问结构,而非单纯规模扩张。

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale

论文配图:The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
图 1 · 摘自论文原文
  • 能力收敛需具备压缩态与可扩展索引双通道的访问结构。
  • 实验证实:64标量状态加全局注意力层后,精确检索误差从0.994降至0.000。
  • 适合关注模型架构设计、信息论边界与可扩展性本质的研究者。

随着模型规模增长,异构网络的表征趋于统一现实模型(柏拉图表征假说)。本文提出其延伸与边界——能力收敛假说(CCH):在固定每标记推理预算下,表征收敛不等于能力收敛。能力将收敛至一类‘访问完备混合体’,即同时具备压缩型O(1)状态通道与可扩展原文索引通道的架构。以无限流中的牛顿苹果问题为见证任务,识别三重资源壁垒:香农壁(禁止任何o(Nb)状态架构)、视野壁(禁止固定窗口)、电路壁(禁止固定深度注意力组合,基于TC0≠NC1假设)。在显式可分性假设下,混合架构支付各壁代价即可突破三壁,表明能力在组合下严格超加性。本文明确区分已证明与推测内容:访问完备性原则基于信息论下界与预注册实验;领域级收敛趋势为经济动机猜想。首次报告预注册小规模测试结果:预测的剪刀差距被测量(精确检索误差0.994→0.000,仅当64标量状态增加一层全局注意力),状态追踪分岔落在注册边界,联合见证显示不可约的双通道解;一预测方向反转,亦如实报告。表征收敛由规模免费提供;能力收敛必须通过访问结构实现。

原文摘要 · Abstract (English)

The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality. We propose its sequel and boundary, the Capability Convergence Hypothesis (CCH): under a fixed per-token inference budget, representational convergence does not entail capability convergence. Capability instead converges toward a class, the access-complete hybrid: any architecture holding both a compressive O(1)-state channel and a scalable verbatim-index channel. We anchor it on a witness task, the Newton's-apple problem in an infinite stream, and name three resource walls: a Shannon wall barring any o(Nb)-state architecture, a horizon wall barring any fixed window, and a circuit wall barring fixed-depth attention-only composition (conditional on TC0 != NC1). Under an explicit separability assumption a hybrid crosses all three by paying each wall's price, so capability is strictly super-additive under composition. We separate what we prove from what we conjecture: the access-completeness principle rests on information-theoretic lower bounds and pre-registered experiments, while the field-level convergence trend is an economics-motivated conjecture. We report the first pre-registered small-scale tests under criteria frozen before the data: the predicted scissors gap is measured (exact-retrieval error 0.994 vs. 0.000 once a 64-scalar state gains one global-attention layer), the state-tracking bifurcation lands at the registered boundary, and a conjunction witness shows an irreducibly two-channel solution; one prediction failed with its direction reversed and is reported as such. Representational convergence is given freely by scale; capability convergence must be purchased by access structure.

模型架构信息论可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。