arXiv:2603.09170cs.ROcs.AI2026-03被引 5

无需遥控数据,直接从人类第一视角视频学做人形机器人全身交互。

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data

  • 用视觉语言模型生成未来动作,再通过跟踪策略控制机器人
  • 在G1人形机器人上实现多样场景感知的自然交互行为
  • 适合想低成本训练人形机器人交互能力的研究者

由于全身遥操作数据成本高昂,实现灵活自然的人形机器人全身交互控制仍具挑战。本文提出ZeroWBC,一种无需遥操作的数据学习框架,仅需配有人类第一视角视频、同步全身运动数据和文本注释即可训练。该方法采用“生成-跟踪”范式解决静态场景下的全身交互控制问题:给定初始第一视角图像和语言指令,经微调的视觉语言模型生成未来人体全身动作标记,解码为连续运动并重定向至人形机器人;随后,参考运动与根部及关键部位轨迹由通用交互运动跟踪策略执行。为提升交互性能,引入面向交互的跟踪奖励,优先保证整体根部与关键部位轨迹对齐,同时保留自然全身动作特征。在Unitree G1人形机器人上的实验表明,ZeroWBC可在无遥控示范的情况下实现多样化场景感知行为,验证了从人类第一视角数据中学习自然人形机器人交互的可扩展性。

原文摘要 · Abstract (English)

Achieving versatile and natural whole-body humanoid interaction control remains challenging due to the high cost of whole-body teleoperation data. We present ZeroWBC, a teleoperation-free framework that learns humanoid whole-body interaction from human egocentric videos paired with synchronized whole-body motion and text annotations. ZeroWBC adopts a generation-then-tracking formulation to tackle the static scene whole-body interaction control problem. Given an initial egocentric image and a language instruction, a fine-tuned Vision-Language Model generates future human whole-body motion tokens, which are decoded into continuous motions and retargeted to the humanoid. The resulting reference motions, together with root and key body-part trajectories, are then executed by a general interactive motion tracking policy. To improve interaction performance, we introduce an interaction-oriented tracking reward that prioritizes global root and key body-part trajectory alignment while preserving natural whole-body motion. Experiments on the Unitree G1 humanoid robot show that ZeroWBC enables diverse scene-aware behaviors without robot teleoperation demonstrations. These results suggest a scalable paradigm for learning natural humanoid whole-body interaction from human egocentric data.

人形机器人视觉语言模型动作生成交互控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。