arXiv:2602.15733cs.ROcs.AI2026-02被引 3

用单目视频训练机器人在复杂地形上自然行走,避免穿模和打滑。

MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction

  • 通过视觉重建人体与环境3D结构,实现动作与地形耦合学习。
  • 在多种复杂地形上实现稳定动态行走,无需昂贵动捕数据。
  • 仅需普通摄像头,适合低成本、大规模机器人训练。

近年来,深度强化学习推动了类人机器人运动控制的重大突破,但高维状态空间与复杂动力学使手动设计动作不切实际,导致对昂贵的运动捕捉(MoCap)数据严重依赖。这些数据常缺乏真实物理环境的几何信息,致使现有运动合成框架中动作与场景脱节,引发接触滑移或网格穿模等物理不一致问题。本文提出 MeshMimic 框架,将3D场景重建与具身智能结合,直接从视频中学习类人机器人与地形的协同交互。利用先进的3D视觉模型,精准分割并重建人体轨迹及地形与物体的三维几何结构;设计基于运动学一致性的优化算法,从噪声视觉重建中提取高质量动作数据,并引入接触无关的重定向方法,将人类-环境交互特征迁移至类人机器人。实验表明,MeshMimic 在多样且具有挑战性的地形上均表现出稳健的高动态性能。该方法证明,仅使用消费级单目传感器即可构建低成本、可扩展的训练管道,为非结构化环境中类人机器人的自主演化提供新路径。

原文摘要 · Abstract (English)

Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-like behaviors. However, the high dimensionality and intricate dynamics of humanoid robots make manual motion design impractical, leading to a heavy reliance on expensive motion capture (MoCap) data. These datasets are not only costly to acquire but also frequently lack the necessary geometric context of the surrounding physical environment. Consequently, existing motion synthesis frameworks often suffer from a decoupling of motion and scene, resulting in physical inconsistencies such as contact slippage or mesh penetration during terrain-aware tasks. In this work, we present MeshMimic, an innovative framework that bridges 3D scene reconstruction and embodied intelligence to enable humanoid robots to learn coupled "motion-terrain" interactions directly from video. By leveraging state-of-the-art 3D vision models, our framework precisely segments and reconstructs both human trajectories and the underlying 3D geometry of terrains and objects. We introduce an optimization algorithm based on kinematic consistency to extract high-quality motion data from noisy visual reconstructions, alongside a contact-invariant retargeting method that transfers human-environment interaction features to the humanoid agent. Experimental results demonstrate that MeshMimic achieves robust, highly dynamic performance across diverse and challenging terrains. Our approach proves that a low-cost pipeline utilizing only consumer-grade monocular sensors can facilitate the training of complex physical interactions, offering a scalable path toward the autonomous evolution of humanoid robots in unstructured environments.

运动控制3D重建机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。