让机器人先看懂视频再模仿行走,实现更自然的运动控制。
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
- 用视觉语言模型解析视频中的动作意图,不依赖姿态重建
- 在第三视角控制下延迟降低80%,任务成功率提升3.7%
- 适合希望实现视频驱动机器人运动的开发者与研究者
人类通过观察视频学习行走,先理解视觉内容再模仿动作。然而当前主流人形机器人运动系统依赖人工标注的动作捕捉数据或稀疏文本指令,导致视觉理解与控制之间存在关键断层。文本到运动的方法存在语义稀疏和流程误差问题,而视频驱动方法仅进行机械姿态模仿,缺乏真正的视觉理解。我们提出RoboMirror,首个无需动作重定向的视频到行走运动框架,贯彻“先理解后模仿”理念。利用视觉语言模型(VLMs),将原始第一/第三人称视频提炼为视觉运动意图,直接作为扩散策略的条件,生成物理合理且语义对齐的运动,无需显式姿态重建或重定向。大量实验验证其有效性:支持通过第一人称视频实现远程操控,第三视角控制延迟降低80%,任务成功率比基线高3.7%。该工作重新定义了人形机器人控制范式,弥合了视觉理解与动作执行之间的鸿沟。
原文摘要 · Abstract (English)
Humans learn locomotion through visual observation, interpreting visual content first before imitating actions. However, state-of-the-art humanoid locomotion systems rely on either curated motion capture trajectories or sparse text commands, leaving a critical gap between visual understanding and control. Text-to-motion methods suffer from semantic sparsity and staged pipeline errors, while video-based approaches only perform mechanical pose mimicry without genuine visual understanding. We propose RoboMirror, the first retargeting-free video-to-locomotion framework embodying "understand before you imitate". Leveraging VLMs, it distills raw egocentric/third-person videos into visual motion intents, which directly condition a diffusion-based policy to generate physically plausible, semantically aligned locomotion without explicit pose reconstruction or retargeting. Extensive experiments validate the effectiveness of RoboMirror, it enables telepresence via egocentric videos, drastically reduces third-person control latency by 80%, and achieves a 3.7% higher task success rate than baselines. By reframing humanoid control around video understanding, we bridge the visual understanding and action gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。