用视觉实现机器人零样本迁移的全身运动操控
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
- 通过人体动作数据训练视觉关键点追踪器,结合任务特定高层策略
- 在仿真中训练的策略可直接部署到真实机器人完成多种操作任务
- 无需外部传感器,适用于室内外复杂环境,适合具身智能研究者
非结构化环境中的人形机器人全身运动操控需要紧密融合第一人称感知与全身控制。现有方法或依赖外部动作捕捉系统,或难以跨任务泛化。我们提出VisualMimic,一种统一第一人称视觉与分层全身控制的视觉模拟到现实框架。该框架结合任务无关的低层关键点追踪器(基于人体动作数据,采用教师-学生训练)与任务相关的高层策略(从视觉和本体感知输入生成关键点指令)。为确保稳定训练,对低层策略注入噪声,并使用人体动作统计量裁剪高层动作。VisualMimic实现了在仿真中训练的视觉运动策略向真实人形机器人的零样本迁移,成功完成箱子搬运、推动、足球带球和踢球等多种运动操控任务。在受控实验室之外,策略也展现出对户外环境的强大泛化能力。视频展示见:https://visualmimic.github.io。
原文摘要 · Abstract (English)
Humanoid loco-manipulation in unstructured environments demands tight integration of egocentric perception and whole-body control. However, existing approaches either depend on external motion capture systems or fail to generalize across diverse tasks. We introduce VisualMimic, a visual sim-to-real framework that unifies egocentric vision with hierarchical whole-body control for humanoid robots. VisualMimic combines a task-agnostic low-level keypoint tracker -- trained from human motion data via a teacher-student scheme -- with a task-specific high-level policy that generates keypoint commands from visual and proprioceptive input. To ensure stable training, we inject noise into the low-level policy and clip high-level actions using human motion statistics. VisualMimic enables zero-shot transfer of visuomotor policies trained in simulation to real humanoid robots, accomplishing a wide range of loco-manipulation tasks such as box lifting, pushing, football dribbling, and kicking. Beyond controlled laboratory settings, our policies also generalize robustly to outdoor environments. Videos are available at: https://visualmimic.github.io .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。