用概率分布融合多视角关键点,提升遮挡下的3D人体姿态估计精度
MEOM: Multi-View Expected-OKS Maximization for Human Pose Triangulation

- 基于热图概率质量对齐设计多视角联合优化目标
- 在无3D标签时达到顶尖模型性能,且推理成本更低
- 适用于遮挡严重场景,适合实际部署的单目多视角系统
传统代数三角测量方法从多视角2D关键点估计3D人体姿态。典型方法通过热图解码2D关键点,但在遮挡下热图可能具有多模态特性,将热图压缩为单一峰值会丢失空间分布信息。本文旨在利用完整热图信息更准确地估计3D姿态,解决两个问题:如何鲁棒融合多视角热图,以及如何评估热图可靠性。为此,提出一种新目标——多视角期望OKS最大化(MEOM),定位各视角概率质量一致的3D关节。同时采用最高密度区域(HDR)校准作为质量诊断,独立于距离度量。该框架涵盖有无3D监督两种设置:无3D标签时,通过最大化MEOM优化预训练热图预测器生成的3D姿态,在模糊的人体3.6M(H36MA)和遮挡的CMU Panoptic数据上表现显著优于现有方法;有3D标签时,端到端训练结合MEOM与MSE损失,在Human3.6M上实现19.11 mm绝对MPJPE,优于当前最优体素方法,且推理成本减半。
原文摘要 · Abstract (English)
Conventional algebraic triangulation solves 3D human pose estimation (HPE) from multi-view 2D keypoints. The typical approach, decoding 2D keypoints from predicted heatmaps, is unreliable as heatmaps can be multimodal under occlusion, and collapsing them into single peaks discards their spatial distribution. We seek to use the entire heatmap to estimate 3D poses more accurately, which requires solving two problems: how to robustly fuse heatmaps across views, and how to assess the reliability of heatmaps. For the former, we introduce a novel objective, Multi-viewExpected-OKS Maximization (MEOM), that locates a 3D joint where the views agree in probability mass. For the latter, we adopt highest-density-region (HDR) calibration as a diagnostic of that mass, independently of distance-based metrics. The proposed framework covers two settings, with and without 3D supervision. Without 3D supervision, we optimize 3D poses from pretrained heatmap predictors by maximizing MEOM, achieving comparable performance with state-of-the-art methods that rely on larger backbones, temporal fusion, and simulated 3D data. On ambiguous Human3.6M (H36MA) and occluded CMU Panoptic frames, the advantage is substantial. When 3D labels are available, we train the model end-to-end with a combined MEOM and MSE loss, achieving 19.11 mm absolute MPJPE on Human3.6M outperforming the state-of-the-art volumetric approach on absolute MPJPE at half the inference cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。