arXiv:2604.03652cs.CV2026-04中稿 · IJCNN 2026, full p…被引 2

通过动态调整时序建模与骨骼约束空间图,实现高效高精度3D人体姿态估计。

Motion-Adaptive Multi-Scale Temporal Modelling with Skeleton-Constrained Spatial Graphs for Efficient 3D Human Pose Estimation

论文配图:Motion-Adaptive Multi-Scale Temporal Modelling with Skeleton-Constrained Spatial Graphs for Efficient 3D Human Pose Estimation
图 1 · 摘自论文原文
  • 动态多尺度时序模块自适应捕捉不同运动节奏
  • 骨骼约束图卷积网络实现关节特异性空间交互建模
  • 在Human3.6M和MPI-INF-3DHP上兼具高精度与低计算开销

从单目视频中精确估计3D人体姿态需要有效建模复杂的时空依赖关系。然而,现有方法在建模时空依赖时往往面临效率与适应性不足的问题,尤其是在密集注意力或固定建模方案下。本文提出MASC-Pose,一种基于骨骼约束空间图的运动自适应多尺度时序建模框架,用于高效3D人体姿态估计。具体而言,该框架引入自适应多尺度时序建模(AMTM)模块,可自适应捕获不同时间尺度下的异质运动动态;同时采用骨骼约束的自适应图卷积网络(SAGCN),实现关节特异性空间交互建模。通过联合实现自适应时序推理与高效空间聚合,该方法在保持高精度的同时具备优异的计算效率。在Human3.6M和MPI-INF-3DHP数据集上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Accurate 3D human pose estimation from monocular videos requires effective modelling of complex spatial and temporal dependencies. However, existing methods often face challenges in efficiency and adaptability when modelling spatial and temporal dependencies, particularly under dense attention or fixed modelling schemes. In this work, we propose MASC-Pose, a Motion-Adaptive multi-scale temporal modelling framework with Skeleton-Constrained spatial graphs for efficient 3D human pose estimation. Specifically, it introduces an Adaptive Multi-scale Temporal Modelling (AMTM) module to adaptively capture heterogeneous motion dynamics at different temporal scales, together with a Skeleton-constrained Adaptive GCN (SAGCN) for joint-specific spatial interaction modelling. By jointly enabling adaptive temporal reasoning and efficient spatial aggregation, our method achieves strong accuracy with high computational efficiency. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets demonstrate the effectiveness of our approach.

3D姿态估计时空建模图神经网络高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。