arXiv:2607.14001cs.LG2026-07
用李雅普诺夫指数做奖励,让强化学习自动找到稳定倒立摆的振荡方法。
Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum
- 以李雅普诺夫指数作为密集奖励信号,引导智能体学习稳定策略。
- 不仅发现经典卡皮察摆现象,还能将摆杆完全稳定在直立位置。
- 适合对物理启发式强化学习感兴趣的科研人员参考。
我们提出将李雅普诺夫特征指数(LCE)作为强化学习中稳定倒立摆(通过垂直运动控制)问题的密集奖励信号。利用LCE,智能体不仅成功发现了著名的卡皮察摆振荡现象,还实现了摆杆摆动的阻尼,最终使其保持在严格竖直的稳定状态。
原文摘要 · Abstract (English)
We suggest using the Lyapunov characteristic exponent (LCE) as a dense reward signal for the reinforcement learning problem of stabilizing the inverted pendulum with vertical motion. With LCE, the agent not only successfully found the oscillatory motion known as the Kapitza pendulum but also damped the pendulum's pivoting, leaving it in a strictly upright position.
强化学习物理启发稳定控制
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。