用可学习的损失函数提升3D人体姿态估计的结构合理性
SEAL-pose: Enhancing 3D Human Pose Estimation via a Learned Loss for Structural Consistency
- 设计基于关节图的可学习损失网络,自动捕捉关节间复杂依赖关系
- 在三个基准上降低关节误差,且不依赖人工规则仍优于有约束模型
- 适合需要高精度姿态估计的场景,如动作分析与虚拟人生成
3D人体姿态估计面临关节间复杂的局部与全局依赖关系。传统监督损失因独立处理每个关节而难以捕捉这些关联。此前研究通过手工设计先验或规则约束来增强结构一致性,但这类方法通常需人工设定且不可微分,无法作为端到端训练目标。本文提出SEAL-pose,一个数据驱动框架:通过可学习的损失网络评估姿态合理性,指导姿态网络训练。不同于手工先验,其基于关节图的设计使损失网络能直接从数据中学习复杂结构依赖。在三个3D HPE基准和八种骨干网络上进行的大量实验表明,SEAL-pose在所有设置下均降低了每关节误差,并提升了姿态合理性。该方法不仅改进了各骨干网络性能,还超越了采用显式结构约束的模型,尽管自身未施加任何约束。最后,分析了损失网络与结构一致性间的关系,并在跨数据集及真实场景下进行了评估。
原文摘要 · Abstract (English)
3D human pose estimation (HPE) is characterized by intricate local and global dependencies among joints. Conventional supervised losses are limited in capturing these correlations because they treat each joint independently. Previous studies have attempted to promote structural consistency through manually designed priors or rule-based constraints; however, these approaches typically require manual specification and are often non-differentiable, limiting their use as end-to-end training objectives. We propose SEAL-pose, a data-driven framework in which a learnable loss-net trains a pose-net by evaluating structural plausibility. Rather than relying on hand-crafted priors, our joint-graph-based design enables the loss-net to learn complex structural dependencies directly from data. Extensive experiments on three 3D HPE benchmarks with eight backbones show that SEAL-pose reduces per-joint errors and improves pose plausibility compared with the corresponding backbones across all settings. Beyond improving each backbone, SEAL-pose also outperforms models with explicit structural constraints, despite not enforcing any such constraints. Finally, we analyze the relationship between the loss-net and structural consistency, and evaluate SEAL-pose in cross-dataset and in-the-wild settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。