arXiv:2605.26879cs.CV2026-05中稿 · CVPR

通过显式建模速度加速度,让单目视频的人体动作更自然

Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos

论文配图:Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos
图 1 · 摘自论文原文
  • 用时序变换器直接预测关节的2D位置、3D速度和加速度
  • 优化后轨迹抖动减少,过平滑现象缓解,动态更真实
  • 适合需要高保真动作还原的研究者和开发者

从单目视频中恢复人体动作常出现过于平滑或动力学不一致的问题,即使关节点位置数值准确。我们发现这源于缺乏可靠的高阶时间信息——速度与加速度,这些是实现真实动量、节奏和高频细节的关键。为此提出HTD-Refine,一个基于显式估计高阶时间动态的后处理框架。核心是PVA-Net,一种时序变换器,可直接从单目视频中推断每个关节的2D位置、3D速度和3D加速度。这些预测的动力学作为软约束,在全局优化中精修世界空间轨迹,显著降低抖动,抑制过平滑,并恢复物理合理的运动。在多个挑战性的野外基准测试上,HTD-Refine持续提升现有先进HMR方法的表现,获得更准确的全局轨迹和更自然的动作动态。结果凸显了高阶时间建模在单目人体动作恢复中的关键作用。

原文摘要 · Abstract (English)

Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate. We observe that this limitation stems from the absence of reliable high-order temporal cues -- velocity and acceleration -- which are essential for reconstructing motion that exhibits realistic momentum, timing, and high-frequency detail. We introduce HTD-Refine, a post-processing framework that augments existing Human Motion Recovery (HMR) pipelines using explicitly estimated high-order temporal dynamics. At the core of our system is PVA-Net, a temporal transformer that infers per-joint 2D positions, 3D velocities, and 3D accelerations directly from a monocular video. These predicted dynamics serve as soft yet informative constraints in a global optimization procedure that refines world-space trajectories, significantly reducing jitter, suppressing over-smoothing, and restoring physically plausible motion. Extensive experiments on challenging in-the-wild benchmarks show that HTD-Refine consistently improves state-of-the-art HMR methods, yielding more accurate global trajectories and substantially more natural motion dynamics. Our results highlight the critical role of high-order temporal modeling in advancing monocular human motion recovery.

动作恢复时序建模单目视频动态优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。