arXiv:2505.07301cs.CV2025-05

用普通视频估计动作,提升3D人体运动预测的泛化能力

Human Motion Prediction via Test-domain-aware Adaptation with Easily-available Human Motions Estimated from Videos

  • 从单目视频中提取2D姿态,转换为模拟动捕的3D动作数据
  • 在新数据上微调模型,使预测更适应真实测试场景
  • 无需昂贵动捕数据,适合实际应用中的动作预测任务

在3D人体运动预测(HMP)中,传统方法依赖昂贵的动作捕捉数据训练模型。然而,这类数据收集成本高,导致数据多样性不足,难以泛化到未见动作或人物。本文提出利用易获取的视频中估计的人体姿态来增强HMP模型。通过我们的处理流程,从单目视频中提取的2D姿态被精确转换为类似动捕风格的3D运动数据。基于这些新增数据进行额外训练,使HMP模型能够适应测试域。实验结果表明,该方法在定量和定性层面均显著提升了模型性能。

原文摘要 · Abstract (English)

In 3D Human Motion Prediction (HMP), conventional methods train HMP models with expensive motion capture data. However, the data collection cost of such motion capture data limits the data diversity, which leads to poor generalizability to unseen motions or subjects. To address this issue, this paper proposes to enhance HMP with additional learning using estimated poses from easily available videos. The 2D poses estimated from the monocular videos are carefully transformed into motion capture-style 3D motions through our pipeline. By additional learning with the obtained motions, the HMP model is adapted to the test domain. The experimental results demonstrate the quantitative and qualitative impact of our method.

人体运动预测视频估计域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。