arXiv:2511.01756cs.CV2025-11

通过频域轨迹一致性提升3D人体姿态估计的时空连贯性

HGFreNet: Hop-hybrid GraphFomer for 3D Human Pose Estimation with Trajectory Consistency in Frequency Domain

  • 设计跳数混合图注意力模块,扩大关节感受野并捕捉全局关联
  • 在频域约束轨迹一致性,显著减少时间序列抖动
  • 适合需要高时序稳定性的动作识别与动画生成任务

单目视频中2D到3D人体姿态恢复是核心挑战,图卷积网络(GCNs)与注意力机制能有效建模关节的时空相关性。然而,2D姿态估计中的深度模糊与误差导致3D轨迹不连贯。现有方法多在时域限制帧间抖动,忽视全局时空关联。为此,本文提出HGFreNet,一种新型GraphFormer架构,结合跳数混合特征聚合与频域轨迹一致性。具体地,设计跳数混合图注意力(HGA)模块,将骨骼关节的k跳邻居分组为混合群,扩大感受野,并利用注意力机制发现全局潜在关联;同时通过频域约束实现全局时间相关性建模。为提供跨帧3D信息以辅助深度推断并保持时序一致性,引入预估网络进行初步3D姿态估计。在Human3.6M和MPI-INF-3DHP两个标准数据集上进行大量实验,结果表明,所提HGFreNet在位置精度与时序一致性方面均优于当前最先进方法。

原文摘要 · Abstract (English)

2D-to-3D human pose lifting is a fundamental challenge for 3D human pose estimation in monocular video, where graph convolutional networks (GCNs) and attention mechanisms have proven to be inherently suitable for encoding the spatial-temporal correlations of skeletal joints. However, depth ambiguity and errors in 2D pose estimation lead to incoherence in the 3D trajectory. Previous studies have attempted to restrict jitters in the time domain, for instance, by constraining the differences between adjacent frames while neglecting the global spatial-temporal correlations of skeletal joint motion. To tackle this problem, we design HGFreNet, a novel GraphFormer architecture with hop-hybrid feature aggregation and 3D trajectory consistency in the frequency domain. Specifically, we propose a hop-hybrid graph attention (HGA) module and a Transformer encoder to model global joint spatial-temporal correlations. The HGA module groups all $k$-hop neighbors of a skeletal joint into a hybrid group to enlarge the receptive field and applies the attention mechanism to discover the latent correlations of these groups globally. We then exploit global temporal correlations by constraining trajectory consistency in the frequency domain. To provide 3D information for depth inference across frames and maintain coherence over time, a preliminary network is applied to estimate the 3D pose. Extensive experiments were conducted on two standard benchmark datasets: Human3.6M and MPI-INF-3DHP. The results demonstrate that the proposed HGFreNet outperforms state-of-the-art (SOTA) methods in terms of positional accuracy and temporal consistency.

3D姿态估计图神经网络频域建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。