用物理最小作用量原理推导出精确反向传播,让梯度计算成为自然过程。
A Physical Theory of Backpropagation: Exact Gradients from the Least-Action Principle
- 将前向过程转为连续时间流,用拉格朗日形式统一建模推理与梯度
- 无需反向传播网络,梯度通过共轭态张力自然产生,结果精确匹配标准BP
- 为神经网络学习提供物理基础,适合研究类脑硬件和力学方法
反向传播通常被视为一种符号化操作:与前向推理拓扑分离,依赖非局部误差信号和全局同步时钟,这些特征在物理现实中无对应。现有物理启发的替代方案仅能近似恢复梯度,需在微扰极限或权重对称条件下成立,与前馈架构不兼容。本文通过哈密顿最小作用量原理,从连续时间前向动力学出发,将非保守系统的拉格朗日形式适配到生成的流上,构建了在双相空间上的统一变分框架。该框架中两个共轭场同时编码激活值与敏感度,单一全局拉格朗日函数控制动态演化:任务损失作为前向流的对称性破缺项,信用分配表现为共轭状态间的张力。因此,推理与梯度计算通过局部相互作用同时展开,无需独立反向电路。最终,标准反向传播被精确还原为该连续流的离散投影。这一视角统一了物理形式与反向传播,为应用经典力学工具(如辛几何、诺特定理、路径积分)分析学习动态开辟了新路径。此外,也为将学习嵌入模拟与类脑硬件提供了理论支持。
原文摘要 · Abstract (English)
Backpropagation is typically presented as a symbolic procedure: a backward pass topologically distinct from inference, with non-local error signals and synchronous global clocking, features with no clear analog in physical reality. Existing physics-inspired alternatives recover gradients only approximately, in vanishing-perturbation limits, or under weight-symmetry constraints incompatible with feedforward architectures. In this paper, we address this gap by deriving exact backpropagation from Hamilton's least-action principle. By recasting the forward dynamics in continuous time and adapting a Lagrangian formalism for non-conservative systems to the resulting flow, we unify inference and gradient computation within a single variational framework on a doubled phase space, whose two conjugate fields jointly encode activations and sensitivities. A single global Lagrangian governs the dynamics: the task loss enters as a symmetry-breaking perturbation of the forward manifold, and credit assignment emerges as the tension that develops between the conjugate states. Inference and gradient computation thus unfold simultaneously through local interactions, requiring no separate backward circuit. Ultimately, standard backpropagation is recovered exactly as the discrete-time projection of this continuous flow. This perspective unifies the formalism of physics with backpropagation, opening a principled pathway for applying tools from classical mechanics - symplectic geometry, Noether's theorem, path-integral methods - to the analysis of learning dynamics. As a downstream consequence, it also points toward analog and neuromorphic substrates in which learning is embodied in the hardware itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。