arXiv:2508.18691cs.RO2025-08被引 3

用人类运动预测模型直接指导机器人动作,无需复杂设计奖励。

Deep Sensorimotor Control by Imitating Predictive Models of Human Motion

  • 用人类运动预测模型零样本迁移指导机器人动作
  • 在多个任务和机器人上表现优于现有方法,无需密集奖励
  • 避免了传统方法中的逆向运动学和对抗损失,更易扩展

随着机器人与人类的具身差距缩小,利用人类与环境交互的数据集为机器人学习提供了新机遇。本文提出一种新方法:通过模仿人类运动预测模型来训练传感器-运动策略。核心洞察是,仿人机器人末端执行器的关键点运动与人体对应关键点运动高度一致,因此可在零样本条件下将人类数据训练的预测模型直接用于机器人。我们训练传感器-运动策略,以跟踪该模型的预测结果,条件为过去机器人状态的历史,并优化稀疏的任务奖励。该方法完全避开基于梯度的运动重定向和对抗损失,充分释放现代人类-场景交互数据集的规模与多样性。实验证明,该方法可跨机器人和任务适用,显著超越现有基线。此外,跟踪人类运动模型可替代精心设计的密集奖励与训练课程。代码、数据与可视化结果见 https://jirl-upenn.github.io/track_reward/

原文摘要 · Abstract (English)

As the embodiment gap between a robot and a human narrows, new opportunities arise to leverage datasets of humans interacting with their surroundings for robot learning. We propose a novel technique for training sensorimotor policies with reinforcement learning by imitating predictive models of human motions. Our key insight is that the motion of keypoints on human-inspired robot end-effectors closely mirrors the motion of corresponding human body keypoints. This enables us to use a model trained to predict future motion on human data \emph{zero-shot} on robot data. We train sensorimotor policies to track the predictions of such a model, conditioned on a history of past robot states, while optimizing a relatively sparse task reward. This approach entirely bypasses gradient-based kinematic retargeting and adversarial losses, which limit existing methods from fully leveraging the scale and diversity of modern human-scene interaction datasets. Empirically, we find that our approach can work across robots and tasks, outperforming existing baselines by a large margin. In addition, we find that tracking a human motion model can substitute for carefully designed dense rewards and curricula in manipulation tasks. Code, data and qualitative results available at https://jirl-upenn.github.io/track_reward/.

机器人控制运动预测零样本迁移强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。