arXiv:2605.00242cs.CVcs.AI2026-05被引 1

用毫米波视频自监督学习人体姿态,提升隐私保护下的精度与鲁棒性

MAEPose: Self-Supervised Spatiotemporal Learning for Human Pose Estimation on mmWave Video

论文配图:MAEPose: Self-Supervised Spatiotemporal Learning for Human Pose Estimation on mmWave Video
图 1 · 摘自论文原文
  • 直接在原始雷达谱图视频上做掩码自编码,保留时空运动信息
  • 在三个数据集上相比顶尖方法误差降低最多22.1%,零样本干扰下误差仅增6.5%
  • 适合关注隐私安全、低计算开销的雷达人体姿态应用

毫米波雷达为人体姿态估计提供了更隐私友好的替代方案。然而,现有方法通常依赖于预提取的中间表示(如稀疏点云或谱图图像),丢弃了雷达视频流中固有的丰富时空信息,且此类信号处理增加了系统复杂度。此外,现有方法多采用端到端监督学习,未利用未标注的原始视频流来学习通用表征。本文提出MAEPose,一种基于掩码自编码的人体姿态估计方法,直接作用于毫米波谱图视频。MAEPose从无标签雷达视频中学习时空运动感知的通用表征,并通过热图解码器实现多帧姿态预测。我们在三个数据集上采用留一人除外交叉验证进行评估,结合严格统计检验。MAEPose在所有设置中均显著优于现有最先进方法,最高可降低22.1%的MPJPE(p<0.05);在零样本旁观者干扰下仍保持鲁棒,误差仅增加6.5%。消融实验表明,预训练和热图解码器均贡献显著,模态分析显示使用范围-多普勒视频输入相比范围-方位或其融合,性能更优且计算成本更低。

原文摘要 · Abstract (English)

Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation. However, existing methods typically rely on pre-extracted intermediate representations such as sparse point clouds or spectrogram images, where the rich spatiotemporal information naturally present in radar video streams is discarded for model learning, while such signal processing adds system complexity. In addition, existing solutions are mainly conducted in an end-to-end supervised manner without leveraging unlabelled raw video streams to learn generalized representations. In this study, we present MAEPose, a masked autoencoding-based human pose estimation approach that operates directly on mmWave spectrogram videos. MAEPose learns spatiotemporal motion-aware generalized representations from unlabelled radar video, and leverages its heatmap decoder for multi-frame pose estimation predictions. We evaluate it across three datasets based on leave-one-person-out cross-validation with rigorous statistical testing. MAEPose consistently outperforms state-of-the-art baselines by up to 22.1% in MPJPE p<0.05, and maintains robust accuracy under zero-shot bystander interference with only a 6.5% error increase. Ablation studies confirm that both the pre-training and the heatmap decoder contribute substantially, while modality analysis indicates that leveraging Range-Doppler video as input achieves better pose estimation performance than Range-Azimuth or their fusion, with lower computational cost.

毫米波雷达姿态估计自监督学习隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。