突破模块化标签局限,从计算过程本质理解神经网络结构差异
Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map
- 以中间状态核空间为视角,定义架构的可区分性边界
- 证明局部与注意力结构在特定前缀下等价但非普遍等价
- 适合研究架构对比与深层推理机制的学者
神经网络架构常通过模块语法、计算图或复合函数来描述,但这些方式回应不同身份问题。本文关注接收端可观测的过程:实际分解形式 $B_j=G_jQ_j$,其中 $Q_j(x)$ 为后续计算提供的中间状态。忽略表现形式而仅保留 $ ext{ker} Q_j$,可得到该切口处保留的前驱区分。对固定分支而言,此扩展影子恰好对无标记满射分解进行分类(至唯一载体重表示),但无法决定标记接收端组织或信息对受限延续的可访问性。组合时相关接口为 $Q_{j,θ}A$,因此下游架构区分依赖上游生成的状态。一个精确的双标记构造表明,局部与注意力模式在一次单射前缀后等价,但在另一前缀后不等价,尽管两种前缀均未丢失前驱信息。结果揭示了模块化标签掩盖的上下文依赖性,推动架构比较回归到被表示过程层面。
原文摘要 · Abstract (English)
Neural architecture is often identified by module syntax, computation graphs, or the composite functions they realize. These descriptions answer different identity questions. We study the represented process available at a receiver: an actual factorization $B_j=G_jQ_j$ in which $Q_j(x)$ is the intermediate state supplied for further computation. Forgetting the presentation and retaining only $\ker Q_j$ yields the predecessor distinctions preserved at that cut. For a fixed branch, this extensional shadow exactly classifies unmarked surjective factorizations up to unique carrier re-presentation, but it does not determine marked receiver organization or the accessibility of retained information to restricted continuations. Under composition the relevant interface is $Q_{j,θ}A$, so downstream architectural distinctions depend on the states produced upstream. An exact two-token construction shows that local and attention schemas are distinction-equivalent after one injective prefix and inequivalent after another, even though neither prefix loses predecessor information. The result exposes a context dependence hidden by modular architecture labels and motivates architecture comparison at the represented-process level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。