arXiv:2608.03234cs.RO2026-08

让机器人运动先验动态适配任务上下文,提升学习效率与自然性。

Learning Context-Aware Motion Priors for Humanoid Control

论文配图:Learning Context-Aware Motion Priors for Humanoid Control
图 1 · 摘自论文原文
  • 基于高奖励轨迹学习上下文与动作的匹配度
  • 在5个任务中提升性能并减少样本需求
  • 无需人工标注或额外技能发现阶段

运动先验能有效指导拟人机器人行为的学习。然而,现有方法通常从完整参考数据集学习通用、无任务特性的先验,并在整个策略训练中统一使用,无法区分当前任务相关的参考动作,可能引入无关或冲突的引导。本文提出上下文感知运动先验(CMP),可在不依赖人工技能标签、数据集划分或独立技能发现的前提下,将通用运动先验自适应地调整至当前任务上下文。具体而言,CMP利用高奖励策略回放学习上下文-动作兼容性,同时通过示范目标保持学习的相关性符合参考分布。由此获得的相关性得分用于重新加权参考监督信号,以训练轻量级上下文条件适配器。我们在对抗运动先验和基于评分匹配的运动先验上验证了该框架,结果表明:在五个拟人机器人控制任务中,CMP持续提升了任务表现与样本效率,实现了有意义的上下文-动作对齐,并对参考分布不平衡具有鲁棒性。这些结果表明,将运动先验适配到任务上下文可为机器人策略学习提供更相关引导。

原文摘要 · Abstract (English)

Motion priors provide powerful guidance for learning naturalistic humanoid behaviors. However, existing methods typically learn a general, task-agnostic prior from the entire reference dataset and apply it uniformly throughout policy training. As a result, the prior cannot distinguish which reference motions are relevant to the current task context, potentially providing irrelevant or conflicting guidance. We present Context-Aware Motion Priors (CMP), a framework that adapts a general motion prior to the current task context without manual skill labels, dataset partitioning, or a separate skill discovery stage. Specifically, CMP learns context-motion compatibility using high-advantage policy rollouts, while a demonstration-based objective keeps the learned relevance grounded in the reference distribution. The resulting relevance scores reweight reference supervision for training a lightweight context-conditioned adapter. To evaluate the effectiveness and generality of CMP, we instantiate it with both Adversarial Motion Priors and Score-Matching Motion Priors. Across five humanoid control tasks, CMP consistently improves task performance and sample efficiency, learns meaningful context-motion alignment, and remains robust to imbalanced reference distributions. These results show that adapting motion priors to task contexts provides more relevant guidance for humanoid policy learning.

机器人控制运动先验上下文感知强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。