解释神经网络如何通过梯度训练发现符号结构。
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
- 将参数升维至测度空间,用Wasserstein流建模训练过程。
- 在群不变性约束下,参数逐渐收缩自由度并解耦优化路径。
- 为神经符号系统设计提供代数与几何结合的理论基础。
我们建立了一个理论框架,解释离散符号结构如何从连续神经网络训练动态中自然涌现。通过将神经参数升维至测度空间,并将训练建模为Wasserstein梯度流,我们证明在几何约束(如群不变性)下,参数测度μ_t会同时经历两个过程:(1) 梯度流分解为若干势函数上的独立优化轨迹;(2) 自由度逐步收缩。这些势函数编码了任务相关的代数约束,在测度空间上的交换半环结构下表现为环同态。随着训练进行,网络从高维探索过渡到符合代数运算的组合表示,自由度降低。我们进一步建立了实现符号任务的数据缩放律,将表征能力与促进符号解的群不变性关联起来。该框架为理解与设计融合连续学习与离散代数推理的神经符号系统提供了严谨基础。
原文摘要 · Abstract (English)
We develop a theoretical framework that explains how discrete symbolic structures can emerge naturally from continuous neural network training dynamics. By lifting neural parameters to a measure space and modeling training as Wasserstein gradient flow, we show that under geometric constraints, such as group invariance, the parameter measure $μ_t$ undergoes two concurrent phenomena: (1) a decoupling of the gradient flow into independent optimization trajectories over some potential functions, and (2) a progressive contraction on the degree of freedom. These potentials encode algebraic constraints relevant to the task and act as ring homomorphisms under a commutative semi-ring structure on the measure space. As training progresses, the network transitions from a high-dimensional exploration to compositional representations that comply with algebraic operations and exhibit a lower degree of freedom. We further establish data scaling laws for realizing symbolic tasks, linking representational capacity to the group invariance that facilitates symbolic solutions. This framework charts a principled foundation for understanding and designing neurosymbolic systems that integrate continuous learning with discrete algebraic reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。