arXiv:2507.09672cs.CV2025-07被引 3

用WiFi信号实现高精度连续人体姿态估计,兼顾隐私与细节动作捕捉。

VST-Pose: A Velocity-Integrated Spatiotem-poral Attention Network for Human WiFi Pose Estimation

  • 双流结构分别建模关节空间关系与时间依赖性
  • 引入速度分支提升对微小动作的敏感度,PCK@50达92.2%
  • 适合智能养老等隐私敏感场景的持续行为分析

基于WiFi的人体姿态估计因其穿透性与隐私优势成为视觉替代方案。本文提出VST-Pose,一种利用WiFi信道状态信息进行精确连续姿态估计的深度学习框架。方法设计了双流结构的ViSTA-Former骨干网络,分别捕获身体关节间的结构关系与时间依赖性。为增强对细微动作的敏感性,引入速度建模分支,学习关键点短期位移模式,提升细粒度运动表征能力。构建专用于智能家居照护场景的2D姿态数据集,实验表明在自建数据集上PCK@50达92.2%,较现有方法提升8.3%。在公开数据集MMFi上的评估验证了模型在3D姿态估计任务中的鲁棒性与有效性。该系统为室内环境中连续人体运动分析提供了可靠且隐私友好的解决方案。代码已开源:https://github.com/CarmenQing/VST-Pose。

原文摘要 · Abstract (English)

WiFi-based human pose estimation has emerged as a promising non-visual alternative approaches due to its pene-trability and privacy advantages. This paper presents VST-Pose, a novel deep learning framework for accurate and continuous pose estimation using WiFi channel state information. The proposed method introduces ViSTA-Former, a spatiotemporal attention backbone with dual-stream architecture that adopts a dual-stream architecture to separately capture temporal dependencies and structural relationships among body joints. To enhance sensitivity to subtle human motions, a velocity modeling branch is integrated into the framework, which learns short-term keypoint dis-placement patterns and improves fine-grained motion representation. We construct a 2D pose dataset specifically designed for smart home care scenarios and demonstrate that our method achieves 92.2% accuracy on the PCK@50 metric, outperforming existing methods by 8.3% in PCK@50 on the self-collected dataset. Further evaluation on the public MMFi dataset confirms the model's robustness and effectiveness in 3D pose estimation tasks. The proposed system provides a reliable and privacy-aware solution for continuous human motion analysis in indoor environments. Our codes are available in https://github.com/CarmenQing/VST-Pose.

WiFi感知姿态估计隐私保护动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。