将毫米波雷达点云转换为人体姿态令牌,提升隐私感知下的姿态估计泛化能力。
Wave2Body: Rethinking mmWave Human Pose Estimation as Radar-to-Body Token Translation

- 用自监督雷达分词器与预训练人体分词器解耦感知与结构学习
- 在M4Human和mmBody数据集上实现更强跨域泛化性能
- 计算开销低,适合实时隐私保护场景应用
毫米波雷达可实现隐私友好的人体感知,但其稀疏点云是视图依赖的电磁反射物理测量,仅间接表征身体姿态。从这种部分且依赖几何的观测中恢复完整3D姿态属于欠约束问题。现有方法直接从雷达-姿态配对数据回归关节坐标,依赖相同有限标注数据同时学习雷达感知、人体结构及二者对齐,易在模糊雷达观测下产生数据集特有捷径。我们提出Wave2Body,一种雷达到人体姿态令牌的翻译框架,通过自监督毫米波分词器、预训练的组合式人体分词器(定义输出空间)以及轻量级翻译器解耦上述学习目标。在M4Human和mmBody数据集上的实验表明,Wave2Body在保持极低训练与推理开销的同时,显著优于以往方法的跨域泛化能力。所有代码与实验结果已公开于https://github.com/Galaxywalk/Wave2Body。
原文摘要 · Abstract (English)
Millimeter-wave (mmWave) radar enables privacy-friendly human sensing, but its sparse point clouds are physical measurements of view-dependent electromagnetic reflections and only indirectly characterize body articulation. Recovering a complete 3D pose from such partial, geometry-dependent observations is therefore under-constrained. Existing methods directly regress joint coordinates from paired radar-pose data, relying on the same limited paired supervision to learn radar perception, human-body structure, and their alignment. This coupling can encourage dataset-specific shortcuts under ambiguous radar observations. We propose Wave2Body, a radar-to-body token translation framework that decouples these learning targets using a self-supervised mmWave tokenizer, a pretrained compositional body tokenizer that defines the output space, and a lightweight translator between them. Experiments on M4Human and mmBody show that Wave2Body achieves stronger cross-domain generalization than previous methods while incurring much lower computational costs for training and inference. All the code and experiment results are publicly available at https://github.com/Galaxywalk/Wave2Body.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。