arXiv:2602.23618cs.CV2026-02

首个带可见性标注的头戴式人体姿态数据集,提升遮挡下姿态估计精度。

Egocentric Visibility-Aware Human Pose Estimation

  • 构建300万帧的头戴式姿态数据集,43.5万帧带关键点可见性标注。
  • 提出EvaPose模型,显式利用可见性信息,显著提升遮挡情况下的估计准确率。
  • 适合研究虚拟现实、增强现实中的遮挡人体姿态估计任务。

使用头戴设备进行自视点人体姿态估计(HPE)对虚拟现实和增强现实应用至关重要,但因关键点被遮挡而面临严峻挑战。现有自视点HPE数据集均未提供关键点可见性标注,且多数方法忽略遮挡问题,对可见与不可见关键点一视同仁,导致可见关键点的预测能力下降。本文首次提出Eva-3M,一个大规模自视点可见性感知的HPE数据集,包含超过300万帧,其中43.5万帧带有关键点可见性标注;同时,我们为现有EMHI数据集补充了可见性标注以促进研究。此外,我们提出EvaPose,一种新型自视点可见性感知的姿态估计方法,显式引入可见性信息以提升估计精度。大量实验验证了真实可见性标签在自视点HPE中的重要价值,并表明EvaPose在Eva-3M和EMHI数据集上均达到当前最优性能。

原文摘要 · Abstract (English)

Egocentric human pose estimation (HPE) using a head-mounted device is crucial for various VR and AR applications, but it faces significant challenges due to keypoint invisibility. Nevertheless, none of the existing egocentric HPE datasets provide keypoint visibility annotations, and the existing methods often overlook the invisibility problem, treating visible and invisible keypoints indiscriminately during estimation. As a result, their capacity to accurately predict visible keypoints is compromised. In this paper, we first present Eva-3M, a large-scale egocentric visibility-aware HPE dataset comprising over 3.0M frames, with 435K of them annotated with keypoint visibility labels. Additionally, we augment the existing EMHI dataset with keypoint visibility annotations to further facilitate the research in this direction. Furthermore, we propose EvaPose, a novel egocentric visibility-aware HPE method that explicitly incorporates visibility information to enhance pose estimation accuracy. Extensive experiments validate the significant value of ground-truth visibility labels in egocentric HPE settings, and demonstrate that our EvaPose achieves state-of-the-art performance in both Eva-3M and EMHI datasets.

姿态估计自视点可见性感知数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。