arXiv:2601.19411cs.ROcs.LG2026-01

让机器人模仿人类动作时,优先保证任务完成,避免错误模仿影响性能。

Task-Centric Policy Optimization from Misaligned Motion Priors

  • 把模仿动作当作条件正则,只在有助于任务时才采纳
  • 在噪声示范下仍保持任务稳定且动作风格一致
  • 适合需要自然动作又强调任务效率的机器人控制场景

人形机器人控制常利用人类示范中的运动先验来生成自然行为。然而,由于身体差异、重定向误差和任务无关变化,这些示范往往次优或与机器人任务不匹配,导致简单模仿会损害任务表现。相反,仅基于任务的强化学习虽能实现最优解,但常产生不自然或不稳定的动作。这暴露了对抗性模仿学习中线性奖励混合的根本缺陷。我们提出任务优先的运动先验(TCMP),将模仿视为条件正则而非平等目标。TCMP在最大化任务改进的同时,仅在模仿信号与任务进展兼容时才引入,实现自适应、几何感知的更新,保留任务可行下降方向,并抑制错误模仿带来的干扰。我们提供了梯度冲突与任务优先不动点的理论分析,并通过人形机器人实验验证:即使在噪声示范下,仍能实现稳健的任务表现与一致的动作风格。

原文摘要 · Abstract (English)

Humanoid control often leverages motion priors from human demonstrations to encourage natural behaviors. However, such demonstrations are frequently suboptimal or misaligned with robotic tasks due to embodiment differences, retargeting errors, and task-irrelevant variations, causing naïve imitation to degrade task performance. Conversely, task-only reinforcement learning admits many task-optimal solutions, often resulting in unnatural or unstable motions. This exposes a fundamental limitation of linear reward mixing in adversarial imitation learning. We propose \emph{Task-Centric Motion Priors} (TCMP), a task-priority adversarial imitation framework that treats imitation as a conditional regularizer rather than a co-equal objective. TCMP maximizes task improvement while incorporating imitation signals only when they are compatible with task progress, yielding an adaptive, geometry-aware update that preserves task-feasible descent and suppresses harmful imitation under misalignment. We provide theoretical analysis of gradient conflict and task-priority stationary points, and validate our claims through humanoid control experiments demonstrating robust task performance with consistent motion style under noisy demonstrations.

机器人控制模仿学习强化学习运动先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。