训练物理振子网络时,记忆、稳定与表达力三者不可兼得。
Between Amnesia and Chaos: A Memory Stability Expressivity Trilemma for Trainable Dissipative Oscillator Networks
- 通过可学习的阻尼参数,端到端训练非线性振子网络的力学属性。
- 梯度稳定性受阻尼限制,训练有效范围随记忆时长收缩至临界点。
- 实验证明:短时记忆下可训练优于冻结,长时记忆则反向,揭示训练边界。
物理储备池计算利用非线性机械动力学,传统上固定底层系统仅训练线性读出,假设底层不可有效训练。本文重新审视此前提,对质量、阻尼和刚度均可学习的非线性振子网络,采用辛积分器进行端到端训练。核心发现为三难困境:记忆时长、梯度稳定性和动态表达力无法同时最大化,三者均由阻尼决定。反向梯度衰减率由阻尼设定,限制了信用传播距离;前向敏感性随最大李雅普诺夫指数指数增长,可用梯度要求阻尼高于稳定阈值。随着阻尼增大,李雅普诺夫指数下降,记忆上限随时长增加而降低,导致稳定训练区域收缩,并在临界点闭合。在二十振子网络上测试,阻尼扫描显示李雅普诺夫指数单调变化且在明确稳定阈值处过零,验证了理论假设。在九个延迟回忆时长下的计算量匹配对比表明,学习底层结构在短时长占优,优势在十一步时长附近消失并反转,符合带宽闭合预测;训练模型自发趋近稳定阈值,探索混沌边缘。分析上限高估实测交叉点约五倍,揭示可观测与可学习梯度间的差距,不加以调参而如实报告。贡献在于确认了何时训练物理底层优于冻结。
原文摘要 · Abstract (English)
Physical reservoir computing harnesses nonlinear mechanical dynamics but, by convention, freezes the substrate and trains only a linear readout, presuming the substrate is not usefully trainable. We revisit that premise for networks of nonlinear oscillators whose mass, damping, and stiffness are learned end-to-end through a symplectic integrator. Our central result is a trilemma: memory horizon, gradient stability, and dynamical expressivity cannot be simultaneously maximized, because all three are governed by the damping. The backward gradient decays at a rate set by the damping, capping how far back credit can propagate, while forward sensitivities grow exponentially in the largest Lyapunov exponent, so usable gradients require damping above a stability floor. Since the Lyapunov exponent falls as damping rises while the memory ceiling falls as the horizon grows, stable training is confined to a band that contracts with horizon and closes at a critical point. We test every step on a twenty-oscillator network. A damping sweep finds the largest Lyapunov exponent monotone and crossing zero at a well-defined stability floor, confirming the theorem's key assumption. A compute-matched comparison of learned versus frozen substrate on delayed recall across nine horizons shows the learned substrate dominating at short horizons and the advantage closing and reversing near a horizon of eleven steps, the predicted signature of band closure; trained models settle near the stability floor, seeking the edge of chaos unprompted. The analytic ceiling overestimates the empirical crossover roughly fivefold, a gap between detectable and learnable gradient that we report rather than tune away. The contribution is a confirmed account of when training a physical substrate beats freezing it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。