arXiv:2510.02566cs.CV2025-10SIGGRAPH被引 8

用视觉直接生成符合物理规律的人形动作,避免传统方法的误差累积。

PhysHMR: Learning Humanoid Control Policies from Vision for Physically Plausible Human Motion Reconstruction

  • 从单目视频直接学习视觉到动作的控制策略,融合全局射线引导与局部视觉特征。
  • 在多个数据集上实现更高视觉精度和物理真实性,优于现有方法。
  • 适合需要真实物理运动的动画、虚拟人生成等场景。

从单目视频重建符合物理规律的人体运动仍是计算机视觉与图形学中的难题。现有方法多依赖基于运动学的姿态估计,因缺乏物理约束而产生不自然结果。以往方法通常采用物理后处理修正,但两阶段设计导致误差累积,限制整体质量。本文提出PhysHMR,一个统一框架,直接在物理模拟器中学习从视觉到人形控制的动作策略,实现兼具物理合理性与视觉一致性的运动重建。关键创新在于像素转射线策略:将2D关键点升维为3D空间射线并映射到全局空间,作为策略输入,提供鲁棒的全局姿态引导,无需依赖噪声较大的3D根节点预测。该软全局引导结合预训练编码器提取的局部视觉特征,使策略能同时推理细节姿态与整体位置。为克服强化学习样本效率低的问题,进一步引入知识蒸馏,将动捕训练专家模型的知识迁移至视觉条件策略,并通过物理激励强化学习进行微调。大量实验表明,PhysHMR在多种场景下均能生成高保真、物理真实的运动,显著优于现有方法,在视觉准确性和物理真实性方面均表现更优。

原文摘要 · Abstract (English)

Reconstructing physically plausible human motion from monocular videos remains a challenging problem in computer vision and graphics. Existing methods primarily focus on kinematics-based pose estimation, often leading to unrealistic results due to the lack of physical constraints. To address such artifacts, prior methods have typically relied on physics-based post-processing following the initial kinematics-based motion estimation. However, this two-stage design introduces error accumulation, ultimately limiting the overall reconstruction quality. In this paper, we present PhysHMR, a unified framework that directly learns a visual-to-action policy for humanoid control in a physics-based simulator, enabling motion reconstruction that is both physically grounded and visually aligned with the input video. A key component of our approach is the pixel-as-ray strategy, which lifts 2D keypoints into 3D spatial rays and transforms them into global space. These rays are incorporated as policy inputs, providing robust global pose guidance without depending on noisy 3D root predictions. This soft global grounding, combined with local visual features from a pretrained encoder, allows the policy to reason over both detailed pose and global positioning. To overcome the sample inefficiency of reinforcement learning, we further introduce a distillation scheme that transfers motion knowledge from a mocap-trained expert to the vision-conditioned policy, which is then refined using physically motivated reinforcement learning rewards. Extensive experiments demonstrate that PhysHMR produces high-fidelity, physically plausible motion across diverse scenarios, outperforming prior approaches in both visual accuracy and physical realism.

动作生成物理模拟视觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。