arXiv:2509.16650eess.SYcs.LG2025-09被引 2

在线学习动态模型,边安全边逼近最优控制

Safe and Near-Optimal Control with Online Dynamics Learning

  • 悲观执行安全策略,乐观探索关键状态
  • 有限时间内以任意精度学习动力学模型
  • 适合高安全性要求的实时控制场景

在未知系统动态下实现最优与安全兼具,是智能体实际部署的核心挑战。本文提出最大安全动态学习概念,在安全策略空间内充分探索。方法在不确定条件下仍以悲观策略确保安全,同时以乐观方式探索有信息量的状态,实现持续在线动态学习。该框架首次实现:在有限时间内将动力学模型学习至任意小容忍度(受噪声影响),全程以高概率保证安全且无需重置。在此基础上,我们提出仅学习至达到近最优性能所需的动态模型的算法。与传统强化学习不同,本方法为非周期在线运行,全程保障安全。在自动驾驶赛车和受空气动力学影响的无人机导航等高风险场景中验证了有效性。

原文摘要 · Abstract (English)

Achieving both optimality and safety under unknown system dynamics is a central challenge in real-world deployment of agents. To address this, we introduce a notion of maximum safe dynamics learning, where sufficient exploration is performed within the space of safe policies. Our method executes $\textit{pessimistically}$ safe policies while $\textit{optimistically}$ exploring informative states and, despite not reaching them due to model uncertainty, ensures continuous online learning of dynamics. The framework achieves first-of-its-kind results: learning the dynamics model sufficiently $-$ up to an arbitrary small tolerance (subject to noise) $-$ in a finite time, while ensuring provably safe operation throughout with high probability and without requiring resets. Building on this, we propose an algorithm to maximize rewards while learning the dynamics $\textit{only to the extent needed}$ to achieve close-to-optimal performance. Unlike typical reinforcement learning (RL) methods, our approach operates online in a non-episodic setting and ensures safety throughout the learning process. We demonstrate the effectiveness of our approach in challenging domains such as autonomous car racing and drone navigation under aerodynamic effects $-$ scenarios where safety is critical and accurate modeling is difficult.

安全控制在线学习强化学习动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。