提出节点中心的时空解耦框架,提升视频人体姿态估计精度。
NCSTR: Node-Centric Decoupled Spatio-Temporal Reasoning for Video-based Human Pose Estimation
- 以关节为中心构建运动感知嵌入,融合外观与动态信息。
- 双分支注意力图建模局部时序传播与全局空间约束,提升一致性。
- 适用于复杂遮挡和模糊场景,适合高精度姿态追踪任务。
基于视频的人体姿态估计仍面临运动模糊、遮挡及复杂时空动态的挑战。现有方法多依赖热图或隐式时空特征聚合,限制了关节拓扑表达能力并削弱跨帧一致性。为此,本文提出一种新型节点中心框架,显式整合视觉、时间与结构推理以实现精准姿态估计。首先设计基于速度的视觉-时序关节嵌入,融合亚像素关节线索与帧间运动,构建兼具外观与运动感知的表征。其次引入注意力驱动的姿态查询编码器,对关节点热图与帧级特征施加注意力,将关节表征映射至姿态感知节点空间,生成图像条件下的关节感知节点嵌入。在此基础上,提出双分支解耦时空注意力图,分别在局部与全局分支中建模时序传播与空间约束推理。最后,设计节点空间专家融合模块,自适应融合两分支互补输出,综合局部与全局线索完成最终关节预测。在三个主流视频姿态基准上的大量实验表明,该方法优于现有最先进方法。结果验证了显式节点中心推理的有效性,为视频人体姿态估计提供了新思路。
原文摘要 · Abstract (English)
Video-based human pose estimation remains challenged by motion blur, occlusion, and complex spatiotemporal dynamics. Existing methods often rely on heatmaps or implicit spatio-temporal feature aggregation, which limits joint topology expressiveness and weakens cross-frame consistency. To address these problems, we propose a novel node-centric framework that explicitly integrates visual, temporal, and structural reasoning for accurate pose estimation. First, we design a visuo-temporal velocity-based joint embedding that fuses sub-pixel joint cues and inter-frame motion to build appearance- and motion-aware representations. Then, we introduce an attention-driven pose-query encoder, which applies attention over joint-wise heatmaps and frame-wise features to map the joint representations into a pose-aware node space, generating image-conditioned joint-aware node embeddings. Building upon these node embeddings, we propose a dual-branch decoupled spatio-temporal attention graph that models temporal propagation and spatial constraint reasoning in specialized local and global branches. Finally, a node-space expert fusion module is proposed to adaptively fuse the complementary outputs from both branches, integrating local and global cues for final joint predictions. Extensive experiments on three widely used video pose benchmarks demonstrate that our method outperforms state-of-the-art methods. The results highlight the value of explicit node-centric reasoning, offering a new perspective for advancing video-based human pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。