arXiv:2512.21573cs.RO2025-12被引 1

用轻量框架从单目视频恢复可直接用于机器人的世界坐标人体动作

World-Coordinate Human Motion Retargeting via SAM 3D Body

  • 以SAM 3D Body为感知主干,用人体骨架表示中间态
  • 通过滑动窗口优化保持动作时序一致性,根节点轨迹更真实
  • 适合需要稳定机器人动作重定向的工业场景应用

从单目视频中恢复具有世界坐标的真人运动并适配至人形机器人,对具身智能和机器人技术至关重要。为避免复杂的SLAM流程或高负载时间模型,我们提出一种轻量化、工程导向的框架:以冻结的SAM 3D Body(3DB)作为感知主干,使用机器人友好的Momentum HumanRig(MHR)表示作为中间表征。方法(i)锁定每位被跟踪主体的身份与骨骼尺度参数,确保骨长在时间上一致;(ii)在低维MHR隐空间中通过高效滑动窗口优化平滑每帧预测;(iii)利用可微分的软足地接触模型与接触感知全局优化,恢复物理合理的全局根轨迹。最后,采用基于运动学感知的两阶段逆运动学管道将重建动作重定向至Unitree G1人形机器人。真实单目视频实验表明,本方法具备稳定的世界轨迹和可靠的机器人重定向效果,说明结构化的人体表示结合轻量物理约束,能从单目输入生成机器人可用的动作。

原文摘要 · Abstract (English)

Recovering world-coordinate human motion from monocular videos with humanoid robot retargeting is significant for embodied intelligence and robotics. To avoid complex SLAM pipelines or heavy temporal models, we propose a lightweight, engineering-oriented framework that leverages SAM 3D Body (3DB) as a frozen perception backbone and uses the Momentum HumanRig (MHR) representation as a robot-friendly intermediate. Our method (i) locks the identity and skeleton-scale parameters of per tracked subject to enforce temporally consistent bone lengths, (ii) smooths per-frame predictions via efficient sliding-window optimization in the low-dimensional MHR latent space, and (iii) recovers physically plausible global root trajectories with a differentiable soft foot-ground contact model and contact-aware global optimization. Finally, we retarget the reconstructed motion to the Unitree G1 humanoid using a kinematics-aware two-stage inverse kinematics pipeline. Results on real monocular videos show that our method has stable world trajectories and reliable robot retargeting, indicating that structured human representations with lightweight physical constraints can yield robot-ready motion from monocular input.

动作重定向单目视频人形机器人轻量建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。