arXiv:2605.01096cs.LGcs.RO2026-05被引 1

无需仿真器,机器人11分钟内实现实时竞速学习

Learning to Race in Minutes: Infoprop Dyna on the Mini Wheelbot

论文配图:Learning to Race in Minutes: Infoprop Dyna on the Mini Wheelbot
图 1 · 摘自论文原文
  • 用不确定性感知的模型强化学习直接从真实交互中训练
  • 微型轮式机器人仅用11分钟完成赛道竞速任务
  • 适合追求高效实机训练的机器人研究者

强化学习有望让具有快速、非线性及不稳定动力学特性的机器人达到性能极限。然而,多数最新进展依赖精心设计的物理仿真器和领域随机化,在合理时间内实现仿真到现实的迁移。本文跳过此类仿真器,证明先进的不确定性感知模型强化学习框架Infoprop Dyna可直接通过真实世界交互实现训练。利用该方法,一款欠驱动单轮机器人Mini Wheelbot仅用11分钟真实经验,便学会在赛道上竞速。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has the potential to enable robots with fast, nonlinear, and unstable dynamics to reach the limits of their performance. However, most recent advances rely on carefully designed physics-based simulators and domain randomization to achieve successful sim-to-real transfer within reasonable wall-clock time. In this work, we bypass the need for such simulators and demonstrate that Infoprop Dyna, a state-of-the-art uncertainty-aware model-based reinforcement learning (MBRL) framework, can enable robots to learn directly from real-world interactions. Using Infoprop Dyna, the Mini Wheelbot, an underactuated unicycle robot, learns to race around a track within 11 minutes of real-world experience.

强化学习机器人控制实时训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。