从物理规律出发,设计出能真正泛化的神经网络架构。
On the Spatiotemporal Dynamics of Generalization in Neural Networks
- 基于局部性、对称性和稳定性构建新型神经元自动机
- 加法任务实现16到100万位数字的100%准确泛化
- 适合研究通用推理与可解释计算系统的学者
为何神经网络无法将16位数相加的规则推广到32位甚至更长?而孩子却能做到?我们提出,这并非工程缺陷,而是违背了计算的物理基本假设。受物理学启发,我们提出三个通用系统必须满足的约束:(1) 局部性——信息以有限速度传播;(2) 对称性——计算法则在时空上不变;(3) 稳定性——系统收敛至抵抗噪声的离散吸引子。由此推导出无需人工设计的时空演化吸引子动态(SEAD)架构:一种通过局部卷积规则迭代直至收敛的神经元自动机。在三项任务中验证理论:(1) 奇偶性任务,通过光锥传播实现完美长度泛化;(2) 加法任务,从L=16到L=10^6实现100%准确推理,且计算自适应输入长度;(3) 规则110,学习图灵完备自动机且无轨迹发散。结果表明,统计学习与逻辑推理之间的鸿沟,不靠参数量堆叠,而需尊重计算的物理本质。
原文摘要 · Abstract (English)
Why do neural networks fail to generalize addition from 16-digit to 32-digit numbers, while a child who learns the rule can apply it to arbitrarily long sequences? We argue that this failure is not an engineering problem but a violation of physical postulates. Drawing inspiration from physics, we identify three constraints that any generalizing system must satisfy: (1) Locality -- information propagates at finite speed; (2) Symmetry -- the laws of computation are invariant across space and time; (3) Stability -- the system converges to discrete attractors that resist noise accumulation. From these postulates, we derive -- rather than design -- the Spatiotemporal Evolution with Attractor Dynamics (SEAD) architecture: a neural cellular automaton where local convolutional rules are iterated until convergence. Experiments on three tasks validate our theory: (1) Parity -- demonstrating perfect length generalization via light-cone propagation; (2) Addition -- achieving scale-invariant inference from L=16 to L=1 million with 100% accuracy, exhibiting input-adaptive computation; (3) Rule 110 -- learning a Turing-complete cellular automaton without trajectory divergence. Our results suggest that the gap between statistical learning and logical reasoning can be bridged -- not by scaling parameters, but by respecting the physics of computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。