用双曲空间建模人体骨骼结构,提升3D姿态估计的几何准确性。
HYPERPOSE: Hyperbolic Kinematic Phase-Space Attention for 3D Human Pose Estimation

- 在双曲空间中构建注意力机制,天然保持骨骼层级结构。
- 在Human3.6M和MPI-INF-3DHP上降低速度误差与体积畸变。
- 适合需要高结构一致性的动作分析任务,如动画生成、医疗评估。
我们提出HYPERPOSE,一种全新的3D人体姿态估计框架,将时空推理完全置于双曲空间$ℝ^d$的洛伦兹模型中,以原生方式保留人体骨骼的分层树状拓扑结构。当前最先进方法依赖变换器和图卷积网络捕捉复杂关节动态,但这些架构仅在欧氏空间运行,与人体固有的树形结构本质不匹配,导致体积指数级畸变且难以维持结构一致性。为此,我们脱离平坦空间,通过双曲运动相空间注意力(HKPSA)原生嵌入复杂关节关系,避免畸变,并引入多尺度窗口化双曲注意力机制,以$O(TW)$复杂度高效建模时序动态。为克服非欧流形训练中的常见不稳定性,HYPERPOSE设计了一套新的黎曼损失函数与不确定性加权课程学习策略,强制施加骨长与速度一致性等物理测地线约束。在Human3.6M和MPI-INF-3DHP数据集上的大量实验表明,该方法在结构与时序一致性方面达到顶尖水平,显著减少体积畸变与速度误差,同时在整体位置精度上建立新基准。
原文摘要 · Abstract (English)
We introduce HYPERPOSE, a novel 3D human pose estimation framework that performs spatio-temporal reasoning entirely within the Lorentz model of hyperbolic space $\mathbb{H}^d$ to natively preserve the hierarchical tree topology of the human skeleton. Current state-of-the-art pose estimators aim to capture complex joint dynamics by relying on transformers and graph convolutional networks. Since these architectures operate exclusively in Euclidean space which fundamentally mismatches the inherent tree structure of the human body, these methods inevitably suffer from exponential volume distortion and struggle to maintain structural coherence. To this end, we depart from flat spaces and aim to improve geometric fidelity with Hyperbolic Kinematic Phase-Space Attention (HKPSA), natively embedding complex joint relationships without distortion, alongside a multi-scale windowed hyperbolic attention mechanism that efficiently models temporal dynamics in $O(TW)$ complexity. Furthermore, to overcome the well-known instability of training non-Euclidean manifolds, HYPERPOSE introduces a novel Riemannian loss suite and an uncertainty-weighted curriculum, enforcing physical geodesic constraints like bone length and velocity consistency. Extensive evaluations on the Human3.6M and MPI-INF-3DHP datasets demonstrate that HYPERPOSE achieves state-of-the-art structural and temporal coherence, significantly reducing both volume distortion and velocity error, while establishing new state-of-the-art benchmarks in overall positional accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。