arXiv:2605.20889cs.CV2026-05中稿 · ICIP 2026, Project…

用单目摄像头+3D地图实现全局人体姿态精准追踪

Map-Mono-Ego: Map-Grounded Global Human Pose Estimation from Monocular Egocentric Video

论文配图:Map-Mono-Ego: Map-Grounded Global Human Pose Estimation from Monocular Egocentric Video
图 1 · 摘自论文原文
  • 结合预扫的3D点云,从单目视频中推算绝对位置
  • 在AIST-Living数据集上超越现有方法,无漂移长期追踪
  • 适合无需外接传感器的日常活动监控场景

单目视角的自我中心人体姿态估计对无处不在的活动监测至关重要。然而,理解用户在环境中的绝对位置仍是挑战。现有方法主要关注相对运动,难以刻画佩戴者在环境中的绝对位置。此外,单目视觉固有的尺度模糊性导致严重平移漂移,限制了无专用多传感器硬件的长期跟踪能力。为此,我们提出MapMonoEgo框架,仅通过单目摄像头与预扫描的3D点云,实现全局一致的人体姿态估计。我们还构建了AIST-Living数据集,将自我中心视频与扫描环境中真实运动标注配对。实验表明,该方法显著优于当前最佳基线,在无需特殊硬件条件下具备实际应用价值。

原文摘要 · Abstract (English)

Monocular egocentric human pose estimation is essential for ubiquitous activity monitoring. However, understanding the user's absolute location within the environment remains a challenge. Existing methods primarily focus on relative motion from an initial position, and tend not to account for the wearer's absolute location within an environment. Furthermore, inherent scale ambiguity in monocular vision leads to severe translational drift, limiting long-term tracking without specialized multi-sensor hardware. To address this, we propose MapMonoEgo, a novel framework achieving globally consistent human pose estimation solely from a monocular camera by leveraging a pre-scanned 3D point cloud. We also introduce AIST-Living dataset, a new dataset pairing egocentric video with ground-truth motion in a scanned environment. Experiments demonstrate that our approach significantly outperforms the state-of-the-art baseline, proving its utility for practical monitoring tasks without specialized hardware.

人体姿态估计单目视觉地图引导长期追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。