arXiv:2506.02754stat.MLcs.LG2025-06NeurIPS被引 2

安全学习随机系统动态,确保训练和部署全程不越界。

Safely Learning Controlled Stochastic Dynamics

  • 用核函数置信区间逐步扩展初始安全控制集,实现安全探索
  • 理论保证安全性和自适应学习率,动态越光滑效率越高
  • 适合机器人、医疗等高风险场景,无需复杂先验

我们解决从离散时间轨迹数据中安全学习受控随机动态的问题,确保系统轨迹在训练和部署期间始终位于预定义的安全区域内。这类安全约束在自动驾驶、金融和生物医学等领域至关重要。本文提出一种方法,通过迭代扩展初始已知安全控制集,利用基于核函数的置信边界,实现安全探索与高效动态估计。训练完成后,所学模型可预测系统动态,并验证任意给定控制的安全性。该方法仅需温和的光滑性假设并依赖一个初始安全控制集,适用于复杂真实系统。我们提供了安全性理论保证,并推导出随真实动态Sobolev正则性提升而加速的自适应学习率。实验表明,该方法在安全性、估计精度和计算效率方面均表现出色。

原文摘要 · Abstract (English)

We address the problem of safely learning controlled stochastic dynamics from discrete-time trajectory observations, ensuring system trajectories remain within predefined safe regions during both training and deployment. Safety-critical constraints of this kind are crucial in applications such as autonomous robotics, finance, and biomedicine. We introduce a method that ensures safe exploration and efficient estimation of system dynamics by iteratively expanding an initial known safe control set using kernel-based confidence bounds. After training, the learned model enables predictions of the system's dynamics and permits safety verification of any given control. Our approach requires only mild smoothness assumptions and access to an initial safe control set, enabling broad applicability to complex real-world systems. We provide theoretical guarantees for safety and derive adaptive learning rates that improve with increasing Sobolev regularity of the true dynamics. Experimental evaluations demonstrate the practical effectiveness of our method in terms of safety, estimation accuracy, and computational efficiency.

安全控制随机系统强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。