提升复杂场景下人体姿态估计,2D与3D相互促进。
BBoxMaskPose v2: Expanding Mutual Conditioning to 3D
- 用概率建模和掩码条件化改进2D姿态估计
- 在COCO上提升1.5 AP,OCHuman上超6 AP并首破50分
- 适合需要高精度多人姿态的复杂场景应用
当前大多数2D人体姿态估计基准已接近饱和,仅在人群密集场景仍有提升空间。本文提出PMPose,一种基于概率框架和掩码条件化的自顶向下2D姿态估计算法,在不牺牲标准场景性能的前提下显著提升密集人群中的姿态估计效果。在此基础上,我们推出BBoxMaskPose v2(BMPv2),融合PMPose与增强版SAM掩码精修模块。BMPv2在COCO数据集上超越现有方法1.5个平均精度(AP)点,在OCHuman上提升6个AP点,首次实现该数据集上超过50 AP的突破。实验表明,利用BMP的2D提示可有效提升3D姿态估计表现,且2D姿态质量的提升直接推动3D结果优化。新构建的OCHuman-Pose数据集结果显示,多人姿态性能更受姿态预测准确率影响,而非检测精度。代码、模型与数据已公开于https://MiraPurkrabek.github.io/BBox-Mask-Pose/。
原文摘要 · Abstract (English)
Most 2D human pose estimation benchmarks are nearly saturated, with the exception of crowded scenes. We introduce PMPose, a top-down 2D pose estimator that incorporates the probabilistic formulation and the mask-conditioning. PMPose improves crowded pose estimation without sacrificing performance on standard scenes. Building on this, we present BBoxMaskPose v2 (BMPv2) integrating PMPose and an enhanced SAM-based mask refinement module. BMPv2 surpasses state-of-the-art by 1.5 average precision (AP) points on COCO and 6 AP points on OCHuman, becoming the first method to exceed 50 AP on OCHuman. We demonstrate that BMP's 2D prompting of 3D model improves 3D pose estimation in crowded scenes and that advances in 2D pose quality directly benefit 3D estimation. Results on the new OCHuman-Pose dataset show that multi-person performance is more affected by pose prediction accuracy than by detection. The code, models, and data are available on https://MiraPurkrabek.github.io/BBox-Mask-Pose/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。