arXiv:2607.16501cs.RO2026-07中稿 · IROS 2026

让机器人安全学习控制,避免保守行为。

Certifiable Safe Model-Based Reinforcement Learning with Control-Affine Dynamics Approximation

论文配图:Certifiable Safe Model-Based Reinforcement Learning with Control-Affine Dynamics Approximation
图 1 · 摘自论文原文
  • 用控制仿射的随机傅里叶特征建模动力学,提升效率
  • 通过自适应置信预测量化安全约束不确定性
  • 适合需要可证明安全性的机器人控制场景

安全模型基于强化学习常结合控制理论与强化学习,使机器人在部分未知系统动力学下安全探索,同时高效生成控制动作。现有方法通常忽略残差模型不确定性中的控制仿射结构(如关于控制输入线性),可能导致行为过于保守或安全保证失效。本文提出一种安全强化学习框架,通过控制仿射随机傅里叶特征(ARFF)学习控制仿射动力学,并结合控制屏障函数(CBF)实现可验证的安全策略。首先,利用ARFF以控制仿射形式建模机器人动力学,计算效率随数据集规模增长,降低模型偏差。其次,采用模型无关、高效的自适应置信预测(ACP)方法,量化由学习的动力学带来的安全约束不确定性。该方法支持基于原则且高效的控制器设计。在小车摆杆和三维四旋翼平台上的仿真结果验证了所提框架的有效性。

原文摘要 · Abstract (English)

Safe model-based reinforcement learning (RL) often bridges control-theoretic analysis and RL for robots to safely explore (partially) unknown system dynamics while deriving control actions for task efficiency. The control performance and safety assurance typically rely on prior knowledge of partially modeled nominal system dynamics and the data-driven models that compensate for residual model uncertainties. However, existing methods often overlook the structure of residual model uncertainties (e.g., components affine in control), which could lead to overly conservative robot behaviors or invalid safety guarantees under the safe learning-based controllers. This paper proposes a safe reinforcement learning framework that learns control-affine dynamics with a certifiable data-driven safe policy using control barrier functions (CBF). Specifically, we first use Control-Affine Random Fourier Features (ARFF) to model robot dynamics in a control-affine form, which offers computational efficiency that scales with dataset size and reduces potential model bias for model-based reinforcement learning. Then, a model-free, efficient uncertainty quantification method using adaptive conformal prediction (ACP) is applied to quantify the uncertainty in the safety constraint arising from the learned control-affine dynamics. This allows for data-driven safety assurance amenable to principled and efficient controller synthesis with CBF. Simulation results on the cartpole and the 3D quadrotor platforms demonstrate the effectiveness of the proposed framework.

安全强化学习控制屏障函数模型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。