构建可扩展的实体智能在线学习系统,支持真实世界多机器人协同训练。
RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI
- 通过硬件抽象层与自适应通信平面统一管理机器人与云边协同训练。
- 支持大规模模型在云端边缘联合训练,实现长时间异步学习。
- 适用于真实场景中多类型机械臂、复杂任务的持续优化,适合工业级应用。
在物理世界中直接进行在线策略学习是具身智能的前沿方向,但受限于无法加速、重置成本高和难以复现等问题,其实质不仅是算法挑战,更是系统工程难题。本文提出USER——一个统一且可扩展的真实世界在线策略学习系统。在系统层面,USER引入硬件抽象层实现机器人统一管理,并设计自适应通信平面以提升云边协同效率;在学习层面,采用全异步训练框架,设计持久化缓存感知回放缓冲区,并提供奖励、算法与策略的可扩展抽象接口。仿真与真实世界实验表明,USER支持多机器人协作、异构机械臂操作、大模型云边联合训练及长期异步训练。这些能力共同构建了真实世界在线策略学习的统一系统基础。
原文摘要 · Abstract (English)
Online policy learning directly in the physical world is a promising yet challenging direction for embodied intelligence. Unlike simulation, real-world systems cannot be arbitrarily accelerated, cheaply reset, or massively replicated, suggesting that real-world policy learning is not merely an algorithmic problem, but inherently a systems problem. We present USER, a \underline{U}nified and extensible \underline{S}yst\underline{E}m for real-world online policy lea\underline{R}ning. On the systems side, USER introduces a hardware abstraction layer for unified robot management and an adaptive communication plane that enables efficient cloud-edge training. On the learning side, USER adopts a fully asynchronous training framework, designs a persistent and cache-aware replay buffer, and provides extensible abstractions for rewards, algorithms, and policies. Experiments in both simulation and the real world demonstrate that USER supports multi-robot coordination, heterogeneous manipulators, cloud-edge training with large models, and long-running asynchronous training. Together, these capabilities establish USER as a unified and extensible systems foundation for real-world online policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。