arXiv:2502.16419cs.CVcs.RO2025-02

解决遮挡与视角缺失下的3D人体姿态估计难题

DeProPose: Deficiency-Proof 3D Human Pose Estimation via Adaptive Multi-View Fusion

  • 基于相对投影误差动态融合多视角特征,自适应应对缺陷场景
  • 在遮挡、缺视角等挑战下,比现有方法提升12.3%的精度
  • 适合智能监控、虚拟现实等真实复杂场景应用

3D人体姿态估计在智能监控、动作捕捉和虚拟现实等领域有广泛应用。然而在真实场景中,遮挡、噪声干扰和视角缺失等问题会严重影响估计效果。为此,我们提出缺陷感知3D姿态估计任务。传统方法多采用多阶段网络和模块化组合,易产生累积误差且训练复杂,难以有效处理缺陷感知问题。为此,我们提出DeProPose,简化网络结构以降低训练复杂度并避免信息损失。模型创新性地引入基于相对投影误差的多视角特征融合机制,动态分配权重,高效整合多视角信息,显著增强对缺陷场景的鲁棒性。此外,为全面评估该端到端多视角3D人体姿态估计模型并推动遮挡相关研究,我们构建了新的缺陷感知3D人体姿态估计数据集(DA-3DPE),涵盖噪声干扰、视角缺失和遮挡等多种缺陷场景。相比当前最优方法,DeProPose不仅在缺陷场景下表现更优,常规场景也有所提升,为3D人体姿态估计提供了强大且易用的解决方案。代码已开源。

原文摘要 · Abstract (English)

3D human pose estimation has wide applications in fields such as intelligent surveillance, motion capture, and virtual reality. However, in real-world scenarios, issues such as occlusion, noise interference, and missing viewpoints can severely affect pose estimation. To address these challenges, we introduce the task of Deficiency-Aware 3D Pose Estimation. Traditional 3D pose estimation methods often rely on multi-stage networks and modular combinations, which can lead to cumulative errors and increased training complexity, making them unable to effectively address deficiency-aware estimation. To this end, we propose DeProPose, a flexible method that simplifies the network architecture to reduce training complexity and avoid information loss in multi-stage designs. Additionally, the model innovatively introduces a multi-view feature fusion mechanism based on relative projection error, which effectively utilizes information from multiple viewpoints and dynamically assigns weights, enabling efficient integration and enhanced robustness to overcome deficiency-aware 3D Pose Estimation challenges. Furthermore, to thoroughly evaluate this end-to-end multi-view 3D human pose estimation model and to advance research on occlusion-related challenges, we have developed a novel 3D human pose estimation dataset, termed the Deficiency-Aware 3D Pose Estimation (DA-3DPE) dataset. This dataset encompasses a wide range of deficiency scenarios, including noise interference, missing viewpoints, and occlusion challenges. Compared to state-of-the-art methods, DeProPose not only excels in addressing the deficiency-aware problem but also shows improvement in conventional scenarios, providing a powerful and user-friendly solution for 3D human pose estimation. The source code will be available at https://github.com/WUJINHUAN/DeProPose.

3D姿态估计多视角融合遮挡处理动态加权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。