让机器人从不完美示范中学习风格,同时保证任务表现最优。
Constrained Style Learning from Imperfect Demonstrations under Task Optimality
- 将问题建模为约束马尔可夫决策过程,用可调拉格朗日乘子控制风格模仿强度。
- 在ANYmal-D上实现14.5%机械能降低和更敏捷步态,任务性能与风格兼得。
- 适合需要自然动作风格但示范不完美的真实机器人应用。
示范学习在机器人领域已证明有效,可用于获取自然的行为模式,如风格化动作和类人灵活性,尤其在难以明确定义风格导向奖励函数时。合成现实任务中的风格化动作通常需平衡任务表现与模仿质量。现有方法通常依赖与任务目标高度对齐的专家示范,但实际示范常不完整或不切实际,导致现有方法在提升风格的同时牺牲任务表现。为此,我们提出将问题建模为约束马尔可夫决策过程(CMDP),在保持近似最优任务表现的前提下优化风格模仿目标。我们引入自适应可调拉格朗日乘子,引导智能体选择性模仿示范,捕捉风格细节而不影响任务性能。我们在多个机器人平台和任务上验证了该方法,均展现出稳健的任务表现与高保真风格学习。在ANYmal-D硬件上,我们实现了14.5%的机械能下降和更敏捷的步态模式,展现了实际应用价值。
原文摘要 · Abstract (English)
Learning from demonstration has proven effective in robotics for acquiring natural behaviors, such as stylistic motions and lifelike agility, particularly when explicitly defining style-oriented reward functions is challenging. Synthesizing stylistic motions for real-world tasks usually requires balancing task performance and imitation quality. Existing methods generally depend on expert demonstrations closely aligned with task objectives. However, practical demonstrations are often incomplete or unrealistic, causing current methods to boost style at the expense of task performance. To address this issue, we propose formulating the problem as a constrained Markov Decision Process (CMDP). Specifically, we optimize a style-imitation objective with constraints to maintain near-optimal task performance. We introduce an adaptively adjustable Lagrangian multiplier to guide the agent to imitate demonstrations selectively, capturing stylistic nuances without compromising task performance. We validate our approach across multiple robotic platforms and tasks, demonstrating both robust task performance and high-fidelity style learning. On ANYmal-D hardware we show a 14.5% drop in mechanical energy and a more agile gait pattern, showcasing real-world benefits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。