用自适应共形预测实现高效安全强化学习,兼顾实时性与不确定性量化。
Computationally and Sample Efficient Safe Reinforcement Learning Using Adaptive Conformal Prediction
- 基于QFF的高斯过程近似加速动态建模,提升计算效率。
- 结合ACP与控制屏障函数,实现在线不确定性量化与安全约束生成。
- 适合需在真实场景中保证安全性的自主系统研发人员使用。
安全是学习型自主系统在现实场景部署中的关键挑战。如何准确量化未知模型的不确定性,以生成可证明安全的控制策略,并促进信息数据的采集,从而实现安全且最优的策略至关重要。此外,数据驱动模型的选择会显著影响实时执行性能和不确定性量化过程。本文提出一种可证明样本高效的周期性安全学习框架,在多种模型选择下仍保持鲁棒性,并具备可量化的不确定性,适用于在线控制任务。首先,采用四阶傅里叶特征(QFF)对高斯过程(GP)的核函数进行近似,实现对未知动态的高效逼近;随后,利用自适应共形预测(ACP)从在线观测中量化不确定性,并与控制屏障函数(CBF)结合,表征基于学习动态的安全控制约束;最后,将基于乐观探索的策略与基于ACP的CBF集成,实现安全探索与近似最优的非线性安全控制。理论证明与仿真结果验证了该框架的有效性与高效性。
原文摘要 · Abstract (English)
Safety is a critical concern in learning-enabled autonomous systems especially when deploying these systems in real-world scenarios. An important challenge is accurately quantifying the uncertainty of unknown models to generate provably safe control policies that facilitate the gathering of informative data, thereby achieving both safe and optimal policies. Additionally, the selection of the data-driven model can significantly impact both the real-time implementation and the uncertainty quantification process. In this paper, we propose a provably sample efficient episodic safe learning framework that remains robust across various model choices with quantified uncertainty for online control tasks. Specifically, we first employ Quadrature Fourier Features (QFF) for kernel function approximation of Gaussian Processes (GPs) to enable efficient approximation of unknown dynamics. Then the Adaptive Conformal Prediction (ACP) is used to quantify the uncertainty from online observations and combined with the Control Barrier Functions (CBF) to characterize the uncertainty-aware safe control constraints under learned dynamics. Finally, an optimism-based exploration strategy is integrated with ACP-based CBFs for safe exploration and near-optimal safe nonlinear control. Theoretical proofs and simulation results are provided to demonstrate the effectiveness and efficiency of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。