首次为非线性系统控制提供对数后悔率,适合高风险场景
Logarithmic Regret for Nonlinear Control
- 基于最优策略持续激励条件,实现对数后悔的快速学习
- 在非持续激励时,后悔率随交互次数平方根增长
- 适用于机器人、医疗等高风险控制任务
我们研究通过序列交互学习控制未知非线性动力系统的问题。针对机器人、医疗等高风险应用中错误代价巨大的情况,探讨快速序列学习的可能性——即学习者相对于全知基线可实现对数后悔。结果表明,在系统动态依赖于未知参数且最优控制策略持续激励的前提下,这类快速学习是可行的。对于非持续激励情形,我们推导出后悔率随交互次数平方根增长的界。这是首个针对依赖未知参数的非线性动力系统控制的后悔边界。理论预测在简单动力系统仿真中得到验证。
原文摘要 · Abstract (English)
We address the problem of learning to control an unknown nonlinear dynamical system through sequential interactions. Motivated by high-stakes applications in which mistakes can be catastrophic, such as robotics and healthcare, we study situations where it is possible for fast sequential learning to occur. Fast sequential learning is characterized by the ability of the learning agent to incur logarithmic regret relative to a fully-informed baseline. We demonstrate that fast sequential learning is achievable in a diverse class of continuous control problems where the system dynamics depend smoothly on unknown parameters, provided the optimal control policy is persistently exciting. Additionally, we derive a regret bound which grows with the square root of the number of interactions for cases where the optimal policy is not persistently exciting. Our results provide the first regret bounds for controlling nonlinear dynamical systems depending nonlinearly on unknown parameters. We validate the trends our theory predicts in simulation on a simple dynamical system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。