构建高真实感机器人持续强化学习基准,支持多种任务与传感器配置。
CRoSS: A Continual Robotic Simulation Suite for Scalable Reinforcement Learning with High Task Diversity and Realistic Physics Simulation
- 基于Gazebo模拟器设计双平台机器人任务,包含视觉/结构参数变化的任务多样性。
- 提供无需物理仿真但可快速运行的纯运动学变体,速度提升100倍。
- 适配持续学习研究者,支持任意传感器与可复现实验环境部署。
持续强化学习(CRL)要求智能体在不遗忘已有策略的前提下,持续学习新任务。本文提出一个基于真实物理模拟的机器人持续强化学习基准套件——CRoSS,使用Gazebo模拟器实现。包含两种机器人平台:一种为带激光雷达、摄像头和碰撞传感器的两轮差速驱动机器人,用于路径追踪与物体推动任务,通过视觉与结构参数变化生成大量不同任务;另一种为七自由度机械臂,在高阶笛卡尔末端位姿控制(模仿Continual World基准)与低阶关节角控制下完成目标到达任务。针对机械臂任务,还提供无需物理仿真的运动学变体,若不涉及传感器读数,可提速两个数量级。CRoSS支持高度可扩展性,能灵活配置任意模拟传感器,确保高物理真实性。为保障可复现性与易用性,提供容器化部署方案(Apptainer),开箱即用,并报告了DQN与策略梯度等标准算法性能,证明其作为可扩展、可复现的CRL研究基准的适用性。
原文摘要 · Abstract (English)
Continual reinforcement learning (CRL) requires agents to learn from a sequence of tasks without forgetting previously acquired policies. In this work, we introduce a novel benchmark suite for CRL based on realistically simulated robots in the Gazebo simulator. Our Continual Robotic Simulation Suite (CRoSS) benchmarks rely on two robotic platforms: a two-wheeled differential-drive robot with lidar, camera and bumper sensor, and a robotic arm with seven joints. The former represent an agent in line-following and object-pushing scenarios, where variation of visual and structural parameters yields a large number of distinct tasks, whereas the latter is used in two goal-reaching scenarios with high-level cartesian hand position control (modeled after the Continual World benchmark), and low-level control based on joint angles. For the robotic arm benchmarks, we provide additional kinematics-only variants that bypass the need for physical simulation (as long as no sensor readings are required), and which can be run two orders of magnitude faster. CRoSS is designed to be easily extensible and enables controlled studies of continual reinforcement learning in robotic settings with high physical realism, and in particular allow the use of almost arbitrary simulated sensors. To ensure reproducibility and ease of use, we provide a containerized setup (Apptainer) that runs out-of-the-box, and report performances of standard RL algorithms, including Deep Q-Networks (DQN) and policy gradient methods. This highlights the suitability as a scalable and reproducible benchmark for CRL research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。