加装后置摄像头可显著提升头戴设备的全身三维姿态估计精度
Bring Your Rear Cameras for Egocentric 3D Human Pose Estimation
- 用Transformer融合前后视角热图与不确定性信息,优化2D关节检测
- 新方法在MPJPE指标上超越当前最优方案超10%
- 适用于需要精准全身追踪的AR/VR、运动分析等场景
以头戴设备(HMD)为载体的自指视角三维人体姿态估计近年来备受关注。尽管前向摄像头对抓手等任务最为理想,但其在全身追踪中因自我遮挡和视场受限而表现不佳,尤其在用户抬头时常见动作下常失效。现有设计普遍忽略身体后方,而这一区域实则蕴含关键3D重建线索。本文首次系统研究后置摄像头的价值,发现仅简单叠加前后视角无法发挥优势,因现有方法依赖独立2D关节检测器且缺乏有效多视角融合机制。为此,提出一种新型基于Transformer的方法,利用多视角信息与热图不确定性来精炼2D关节估计,从而提升3D姿态追踪性能。同时构建两个大规模新数据集Ego4View-Syn与Ego4View-RW用于后视评估。实验表明,引入后置视角显著优于仅前端配置,所提方法在MPJPE指标上相较现有最优方法提升超过10%。代码、模型与数据集已公开。
原文摘要 · Abstract (English)
Egocentric 3D human pose estimation has been actively studied using cameras installed in front of a head-mounted device (HMD). While frontal placement is the optimal and the only option for some tasks, such as hand tracking, it remains unclear if the same holds for full-body tracking due to self-occlusion and limited field-of-view coverage. Notably, even the state-of-the-art methods often fail to estimate accurate 3D poses in many scenarios, such as when HMD users tilt their heads upward -- a common motion in human activities. A key limitation of existing HMD designs is their neglect of the back of the body, despite its potential to provide crucial 3D reconstruction cues. Hence, this paper investigates the usefulness of rear cameras for full-body tracking. We also show that simply adding rear views to the frontal inputs is not optimal for existing methods due to their dependence on individual 2D joint detectors without effective multi-view integration. To address this issue, we propose a new transformer-based method that refines 2D joint heatmap estimation with multi-view information and heatmap uncertainty, thereby improving 3D pose tracking. Also, we introduce two new large-scale datasets, Ego4View-Syn and Ego4View-RW, for a rear-view evaluation. Our experiments show that the new camera configurations with back views provide superior support for 3D pose tracking compared to only frontal placements. The proposed method achieves significant improvement over the current state of the art (>10% on MPJPE). The source code, trained models, and datasets are available on our project page at https://4dqv.mpi-inf.mpg.de/EgoRear/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。