让机器人学会连贯完成复杂篮球动作,无需依赖球的轨迹信息。
Learning to Ball: Composing Policies for Long-Horizon Basketball Moves
- 设计新框架实现不同运动技能在长时序任务中的无缝组合。
- 在模拟环境中成功控制角色完成真实用户指令下的连续篮球动作。
- 适合需要多阶段协同控制的机器人动作生成场景。
学习多阶段、长时序任务(如篮球动作)的控制策略仍对强化学习构成挑战,主要源于策略间需平滑衔接与技能转换。长时序任务通常由具有明确目标的子任务和目标不明确但关键的过渡子任务组成。现有方法如专家混合与技能串联,在个体策略无显著共享状态或缺乏清晰起始/终止状态时表现不佳。本文提出一种新颖的策略整合框架,支持在中间状态定义模糊的情况下,组合差异较大的运动技能。在此基础上,引入高层软路由机制,实现子任务间平滑且鲁棒的切换。我们在一系列基础篮球技能及复杂过渡上评估该框架。通过该方法训练的策略能有效控制模拟角色与球互动,完成由实时用户指令指定的长时序任务,且无需依赖球的轨迹参考。
原文摘要 · Abstract (English)
Learning a control policy for a multi-phase, long-horizon task, such as basketball maneuvers, remains challenging for reinforcement learning approaches due to the need for seamless policy composition and transitions between skills. A long-horizon task typically consists of distinct subtasks with well-defined goals, separated by transitional subtasks with unclear goals but critical to the success of the entire task. Existing methods like the mixture of experts and skill chaining struggle with tasks where individual policies do not share significant commonly explored states or lack well-defined initial and terminal states between different phases. In this paper, we introduce a novel policy integration framework to enable the composition of drastically different motor skills in multi-phase long-horizon tasks with ill-defined intermediate states. Based on that, we further introduce a high-level soft router to enable seamless and robust transitions between the subtasks. We evaluate our framework on a set of fundamental basketball skills and challenging transitions. Policies trained by our approach can effectively control the simulated character to interact with the ball and accomplish the long-horizon task specified by real-time user commands, without relying on ball trajectory references.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。