用轨迹线指导强化学习,让自行车机器人学会多种高难度特技动作
LineRides: Line-Guided Reinforcement Learning for Bicycle Robot Stunts

- 通过用户提供的空间轨迹线和关键朝向,无需示范即可训练动作
- 支持5种特技动作,可实现正常骑行与特技间的无缝切换
- 适合开发新型机器人特技行为,尤其缺乏参考动作的场景
在强化学习中设计敏捷机器人动作的奖励函数仍具挑战性,而基于示范的方法通常需要不可获取的参考运动。我们提出LineRides,一种基于轨迹线的引导学习框架,使定制自行车机器人能从用户提供的空间路径和稀疏关键朝向中,无需示范或显式时间信息,自主习得多样且可命令的特技行为。该方法通过跟踪容差处理物理不可行路径,利用沿路径行进距离衡量进展以解决时间歧义,并借助位置与序列双重关键朝向明确动作细节。我们在Ultra Mobility Vehicle(UMV)上评估了LineRides,结果表明,采用本方法训练的策略支持正常驾驶与特技执行之间的无缝过渡,可实现五种不同特技:MiniHop、LargeHop、ThreePointTurn、Backflip 和 DriftTurn。
原文摘要 · Abstract (English)
Designing reward functions for agile robotic maneuvers in reinforcement learning remains difficult, and demonstration-based approaches often require reference motions that are unavailable for novel platforms or extreme stunts. We present LineRides, a line-guided learning framework that enables a custom bicycle robot to acquire diverse, commandable stunt behaviors from a user-provided spatial guideline and sparse key-orientations, without demonstrations or explicit timing. LineRides handles physically infeasible guidelines using a tracking margin that permits controlled deviation, resolves temporal ambiguity by measuring progress via traveled distance along the guideline, and disambiguates motion details through position- and sequence-based key-orientations. We evaluate LineRides on the Ultra Mobility Vehicle (UMV) and show that the policy trained with our methods supports seamless transitions between normal driving and stunt execution, enabling five distinct stunts on command: MiniHop, LargeHop, ThreePointTurn, Backflip, and DriftTurn.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。