arXiv:2511.00112cs.ROcs.AI2025-11

让机器人在真实环境中安全学习高性能控制策略。

Real-DRL: Teach and Learn in Reality

  • 采用学生-教师协同机制,实时融合物理模型与深度强化学习。
  • 在真实四足机器人上实现安全优先的自主学习,性能优于传统方法。
  • 适合需要高安全性、动态适应的真实世界控制系统开发。

本文提出Real-DRL框架,用于安全关键型自主系统在真实物理系统(real plants)中的运行时学习。该框架包含三个交互组件:DRL-Student(深度强化学习学生)、PHY-Teacher(基于物理模型的教师)和Trigger(触发器)。DRL-Student采用自学习与教学相长双重机制,并结合实时安全感知批采样;PHY-Teacher基于物理模型设计仅关注安全功能的动作策略,实现实时补丁,支持教学相长并保障系统安全;Trigger协调二者交互。该框架有效应对未知未知和Sim2Real差距带来的安全挑战。其显著特性包括:1)保证安全性;2)自动分层学习(先安全后高性能);3)安全感知批采样,缓解极端情况导致的学习经验失衡。在真实四足机器人、NVIDIA Isaac Gym中的四足机器人及倒立摆系统上的实验,结合对比与消融研究,验证了Real-DRL的有效性与独特优势。

原文摘要 · Abstract (English)

This paper introduces the Real-DRL framework for safety-critical autonomous systems, enabling runtime learning of a deep reinforcement learning (DRL) agent to develop safe and high-performance action policies in real plants (i.e., real physical systems to be controlled), while prioritizing safety! The Real-DRL consists of three interactive components: a DRL-Student, a PHY-Teacher, and a Trigger. The DRL-Student is a DRL agent that innovates in the dual self-learning and teaching-to-learn paradigm and the real-time safety-informed batch sampling. On the other hand, PHY-Teacher is a physics-model-based design of action policies that focuses solely on safety-critical functions. PHY-Teacher is novel in its real-time patch for two key missions: i) fostering the teaching-to-learn paradigm for DRL-Student and ii) backing up the safety of real plants. The Trigger manages the interaction between the DRL-Student and the PHY-Teacher. Powered by the three interactive components, the Real-DRL can effectively address safety challenges that arise from the unknown unknowns and the Sim2Real gap. Additionally, Real-DRL notably features i) assured safety, ii) automatic hierarchy learning (i.e., safety-first learning and then high-performance learning), and iii) safety-informed batch sampling to address the learning experience imbalance caused by corner cases. Experiments with a real quadruped robot, a quadruped robot in NVIDIA Isaac Gym, and a cart-pole system, along with comparisons and ablation studies, demonstrate the Real-DRL's effectiveness and unique features.

强化学习机器人控制安全学习真实环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。