arXiv:2507.16850cs.CVcs.AI2025-07

用几何先验实现实时单目人体三维姿态估计

Toward a Real-Time Framework for Accurate Monocular 3D Human Pose Estimation with Geometric Priors

  • 2D关键点检测后结合相机参数与人体结构先验进行3D还原
  • 在无需专用硬件下实现快速、精准且个性化的姿态估计
  • 适合边缘设备部署,兼顾准确性与可解释性

单目三维人体姿态估计在实时场景和非受限环境中仍具挑战性。直接图像到3D的方法依赖大量标注数据和重型模型,而2D到3D提升方法更轻量灵活,尤其在引入先验知识后表现更优。本文提出一种框架,结合实时2D关键点检测与几何感知的2D-to-3D提升,显式利用已知相机内参和个体特定解剖学先验。基于自校准与生物力学约束逆运动学的最新进展,从动作捕捉(MoCap)和合成数据集中生成大规模、合理的2D-3D训练样本对。该方法可在不依赖专用硬件的前提下,实现快速、个性化且高精度的单目三维姿态估计,推动数据驱动学习与模型驱动先验的融合,提升边缘设备上真实世界人体运动捕捉的准确率、可解释性与可部署性。

原文摘要 · Abstract (English)

Monocular 3D human pose estimation remains a challenging and ill-posed problem, particularly in real-time settings and unconstrained environments. While direct imageto-3D approaches require large annotated datasets and heavy models, 2D-to-3D lifting offers a more lightweight and flexible alternative-especially when enhanced with prior knowledge. In this work, we propose a framework that combines real-time 2D keypoint detection with geometry-aware 2D-to-3D lifting, explicitly leveraging known camera intrinsics and subject-specific anatomical priors. Our approach builds on recent advances in self-calibration and biomechanically-constrained inverse kinematics to generate large-scale, plausible 2D-3D training pairs from MoCap and synthetic datasets. We discuss how these ingredients can enable fast, personalized, and accurate 3D pose estimation from monocular images without requiring specialized hardware. This proposal aims to foster discussion on bridging data-driven learning and model-based priors to improve accuracy, interpretability, and deployability of 3D human motion capture on edge devices in the wild.

人体姿态估计单目3D几何先验边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。