arXiv:2505.03729cs.ROcs.CV2025-05被引 107

看视频学动作,让机器人能根据环境自主完成上下楼梯、坐凳等复杂动作。

Visual Imitation Enables Contextual Humanoid Control

  • 通过分析日常视频,重建人与环境,生成全身控制策略。
  • 单个策略实现上下楼梯、坐立等动态动作,真实机器人测试稳定可重复。
  • 无需复杂编程,适合快速部署到多样现实场景的机器人任务。

如何让机器人在真实环境中通过环境上下文完成爬楼梯、坐椅子等任务?最简单的方法是直接示范——随意拍摄一段人类动作视频并输入给机器人。我们提出 VIDEOMIMIC,一个从实到虚再到实的端到端流程:从日常视频中挖掘数据,联合重建人体与环境,并生成适用于类人机器人执行相应技能的全身控制策略。我们在真实类人机器人上验证了该方法,展示了包括上下楼梯、从椅子或长椅起身坐下在内的多种动态全身动作,全部由单一策略实现,且策略受环境状态和全局根节点指令驱动。VIDEOMIMIC为教类人机器人在多样化真实环境中操作提供了可扩展的路径。

原文摘要 · Abstract (English)

How can we teach humanoids to climb staircases and sit on chairs using the surrounding environment context? Arguably, the simplest way is to just show them-casually capture a human motion video and feed it to humanoids. We introduce VIDEOMIMIC, a real-to-sim-to-real pipeline that mines everyday videos, jointly reconstructs the humans and the environment, and produces whole-body control policies for humanoid robots that perform the corresponding skills. We demonstrate the results of our pipeline on real humanoid robots, showing robust, repeatable contextual control such as staircase ascents and descents, sitting and standing from chairs and benches, as well as other dynamic whole-body skills-all from a single policy, conditioned on the environment and global root commands. VIDEOMIMIC offers a scalable path towards teaching humanoids to operate in diverse real-world environments.

视觉模仿类人机器人动作控制真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。