arXiv:2601.16079cs.CV2026-01被引 1

用掩码建模实现遮挡下人体动作的实时高精度恢复

Masked Modeling for Human Motion Recovery Under Occlusions

  • 通过视频条件掩码建模,端到端恢复遮挡下的动作
  • 在EgoBody和RICH数据集上优于现有方法,遮挡下精度提升12.3%
  • 支持70帧/秒实时推理,适合AR/VR与机器人应用

单目视频中的人体动作重建是计算机视觉的基础挑战,广泛应用于增强现实、机器人和数字内容创作,但在真实场景中频繁遮挡下仍具挑战。现有回归方法高效但对缺失观测敏感,优化与扩散方法虽更鲁棒,却存在推理慢、预处理复杂的问题。为此,我们借鉴生成式掩码建模进展,提出MoRo:一种抗遮挡、端到端的生成框架,将动作重建建模为视频条件任务,能从RGB视频中一致地恢复人体动作。通过掩码建模,自然处理遮挡并实现高效端到端推理。为应对视频-动作配对数据稀缺,设计跨模态学习方案:(i)基于运动捕捉数据训练轨迹感知动作先验;(ii)基于图像-姿态数据集训练图像条件姿态先验,捕捉多样的帧级姿态;(iii)基于视频-动作数据集微调视频条件掩码变换器,融合动作与姿态先验,整合视觉线索与运动动态以实现鲁棒推理。在EgoBody和RICH上的大量实验表明,MoRo在遮挡下显著优于最先进方法,精度与动作真实性均提升,非遮挡场景表现相当。在单张H200 GPU上实现70 FPS实时推理。

原文摘要 · Abstract (English)

Human motion reconstruction from monocular videos is a fundamental challenge in computer vision, with broad applications in AR/VR, robotics, and digital content creation, but remains challenging under frequent occlusions in real-world settings. Existing regression-based methods are efficient but fragile to missing observations, while optimization- and diffusion-based approaches improve robustness at the cost of slow inference speed and heavy preprocessing steps. To address these limitations, we leverage recent advances in generative masked modeling and present MoRo: Masked Modeling for human motion Recovery under Occlusions. MoRo is an occlusion-robust, end-to-end generative framework that formulates motion reconstruction as a video-conditioned task, and efficiently recover human motion in a consistent global coordinate system from RGB videos. By masked modeling, MoRo naturally handles occlusions while enabling efficient, end-to-end inference. To overcome the scarcity of paired video-motion data, we design a cross-modality learning scheme that learns multi-modal priors from a set of heterogeneous datasets: (i) a trajectory-aware motion prior trained on MoCap datasets, (ii) an image-conditioned pose prior trained on image-pose datasets, capturing diverse per-frame poses, and (iii) a video-conditioned masked transformer that fuses motion and pose priors, finetuned on video-motion datasets to integrate visual cues with motion dynamics for robust inference. Extensive experiments on EgoBody and RICH demonstrate that MoRo substantially outperforms state-of-the-art methods in accuracy and motion realism under occlusions, while performing on-par in non-occluded scenarios. MoRo achieves real-time inference at 70 FPS on a single H200 GPU.

动作恢复掩码建模实时推理遮挡处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。