用波动散射重新定义反向传播,实现无需锁步的神经网络训练。
Unlocked Backpropagation using Wave Scattering
- 引入优化时间维,将反向传播转为波动散射问题。
- 通过波的反射最小化实现参数更新,可导出梯度下降等算法。
- 物理系统只要支持波散射与耗散,就能自动优化。
机器学习中的反向传播算法与最优控制理论中的最大值原理均表现为两点边值问题,导致“前向-反向”锁定。我们通过引入额外的“优化时间”维度,将最大值原理重构为双曲初值问题。引入具有有限传播速度的反向传播波变量,并将优化问题重新表述为它们之间的散射关系。该问题的松弛形式可被理解为一个物理系统,通过均衡并改变自身物理属性来最小化反射。我们将此连续理论离散化,推导出一类完全解耦的训练算法,适用于神经网络。不同的参数动力学(包括梯度下降)可通过在参数端口要求耗散和反射最小化而获得。这些结果还表明,任何支持波散射与耗散的物理基底,均可被解释为求解优化问题。
原文摘要 · Abstract (English)
Both the backpropagation algorithm in machine learning and the maximum principle in optimal control theory are posed as a two-point boundary problem, resulting in a "forward-backward" lock. We derive a reformulation of the maximum principle in optimal control theory as a hyperbolic initial value problem by introducing an additional "optimization time" dimension. We introduce counter-propagating wave variables with finite propagation speed and recast the optimization problem in terms of scattering relationships between them. This relaxation of the original problem can be interpreted as a physical system that equilibrates and changes its physical properties in order to minimize reflections. We discretize this continuum theory to derive a family of fully unlocked algorithms suitable for training neural networks. Different parameter dynamics, including gradient descent, can be derived by demanding dissipation and minimization of reflections at parameter ports. These results also imply that any physical substrate that supports the scattering and dissipation of waves can be interpreted as solving an optimization problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。