用几何信息消除WiFi布局依赖,实现跨场景人体3D姿态精准估计
Breaking Coordinate Overfitting: Geometry-Aware WiFi Sensing for Cross-Layout 3D Pose Estimation
- 引入几何感知的坐标统一机制,用棋盘和照片对齐视觉与WiFi数据
- 在统一空间中编码设备位置,使模型区分动作与部署布局
- 首个实现跨布局泛化的WiFi姿态估计,适合智能交互与隐私场景
基于WiFi的3D人体姿态估计为智能交互提供了低成本且保护隐私的替代方案。然而,现有方法依赖视觉3D姿态作为监督信号,并直接将信道状态信息(CSI)回归到摄像头坐标系,导致坐标过拟合:模型记忆特定部署下的天线布局而非仅学习与活动相关的表征,造成严重泛化失败。为此,我们提出PerceptAlign,首个面向跨布局的几何条件化框架。该框架通过两个棋盘和少量照片,实现轻量级坐标统一,将WiFi与视觉测量对齐至共享3D空间。在此空间中,将校准的收发器位置编码为高维嵌入,并与CSI特征融合,使模型显式以设备几何为条件变量。这一设计迫使网络解耦人体运动与部署布局,首次实现鲁棒的、布局无关的WiFi姿态估计。为支持系统评估,我们构建了当前最大规模的跨域3D WiFi姿态估计数据集,包含21名受试者、5个场景、18种动作和7种设备布局。实验表明,PerceptAlign相比顶尖基线,在同域误差降低12.3%,跨域误差降低超60%。这些结果确立了几何条件学习在可扩展、实用化WiFi感知中的可行性。
原文摘要 · Abstract (English)
WiFi-based 3D human pose estimation offers a low-cost and privacy-preserving alternative to vision-based systems for smart interaction. However, existing approaches rely on visual 3D poses as supervision and directly regress CSI to a camera-based coordinate system. We find that this practice leads to coordinate overfitting: models memorize deployment-specific WiFi transceiver layouts rather than only learning activity-relevant representations, resulting in severe generalization failures. To address this challenge, we present PerceptAlign, the first geometry-conditioned framework for WiFi-based cross-layout pose estimation. PerceptAlign introduces a lightweight coordinate unification procedure that aligns WiFi and vision measurements in a shared 3D space using only two checkerboards and a few photos. Within this unified space, it encodes calibrated transceiver positions into high-dimensional embeddings and fuses them with CSI features, making the model explicitly aware of device geometry as a conditional variable. This design forces the network to disentangle human motion from deployment layouts, enabling robust and, for the first time, layout-invariant WiFi pose estimation. To support systematic evaluation, we construct the largest cross-domain 3D WiFi pose estimation dataset to date, comprising 21 subjects, 5 scenes, 18 actions, and 7 device layouts. Experiments show that PerceptAlign reduces in-domain error by 12.3% and cross-domain error by more than 60% compared to state-of-the-art baselines. These results establish geometry-conditioned learning as a viable path toward scalable and practical WiFi sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。