arXiv:2510.18518cs.RO2025-10

用在线优化建模,让机器人在真实世界高效自学习控制策略。

Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization

  • 基于实时交互数据在线构建动力学模型,指导策略更新。
  • 样本效率显著提升,数小时内达到与模型无关方法相当性能。
  • 适合需快速适应动态变化的真实机器人控制任务。

我们提出一种适用于复杂机器人系统直接在真实世界中控制的在线模型基于强化学习算法。与依赖大量离线仿真和无模型策略优化的典型模拟到现实流程不同,本方法利用实时交互数据构建动力学模型,并基于所学模型进行策略更新。该高效模型基强化学习方案显著减少训练控制策略所需的样本数,支持直接在真实世界滚动数据上训练,大幅降低模拟数据偏差的影响,并促进高性能控制策略的搜索。通过在线优化分析,在随机在线优化假设下推导出次线性遗憾界,为随着交互数据增加性能持续提升提供形式化保证。在液压挖掘机臂和软体机械臂上的实验表明,相比模型无关强化学习方法,该算法展现出强大样本效率,数小时内即达相当性能;当负载条件随机变化时,也观察到稳健的动态适应能力。本方法为一系列复杂控制任务的高效可靠机载学习铺平道路。

原文摘要 · Abstract (English)

We present an online model-based reinforcement learning algorithm suitable for controlling complex robotic systems directly in the real world. Unlike prevailing sim-to-real pipelines that rely on extensive offline simulation and model-free policy optimization, our method builds a dynamics model from real-time interaction data and performs policy updates guided by the learned dynamics model. This efficient model-based reinforcement learning scheme significantly reduces the number of samples to train control policies, enabling direct training on real-world rollout data. This significantly reduces the influence of bias in the simulated data, and facilitates the search for high-performance control policies. We adopt online optimization analysis to derive sublinear regret bounds under stochastic online optimization assumptions, providing formal guarantees on performance improvement as more interaction data are collected. Experimental evaluations were performed on a hydraulic excavator arm and a soft robot arm, where the algorithm demonstrates strong sample efficiency compared to model-free reinforcement learning methods, reaching comparable performance within hours. Robust adaptation to shifting dynamics was also observed when the payload condition was randomized. Our approach paves the way toward efficient and reliable on-robot learning for a broad class of challenging control tasks.

机器人控制强化学习在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。