arXiv:2511.18792cs.CVcs.IT2025-11被引 2

用大规模多样数据训练模型,提升无线感知跨场景泛化能力。

Scale What Counts, Mask What Matters: Evaluating Foundation Models for Zero-Shot Cross-Domain Wi-Fi Sensing

  • 基于14个数据集的130万条信号数据,采用掩码自编码预训练。
  • 数据量越大,跨域性能提升越明显,模型规模提升收益有限。
  • 在动作、手势识别任务中,准确率提升2.2%至15.7%,适合部署应用。

尽管无线感知提供了隐私友好的替代方案,但其实际应用受限于跨领域鲁棒性差的问题。现有模型在不同环境、设备或用户间难以泛化,而这一“领域偏移”问题因数据量小且分散而加剧。本文突破传统范式,采用基础模型方法,利用最大且最异构的无线信道状态信息(Wi-Fi CSI)数据集进行掩码自编码(MAE)风格预训练。研究覆盖超过130万条样本,来自14个数据集,涵盖4种设备、2.4/5/6 GHz频段及20至160 MHz带宽。首次系统性地分离数据多样性与模型容量对跨域性能的影响。实验表明:预训练数据量增加时,未见领域性能呈对数线性提升,说明数据规模与多样性是关键;当前数据量下,更大模型仅带来微弱增益,表明数据而非模型容量是瓶颈。在人体动作识别、手势识别和用户识别任务中,大规模预训练使跨域准确率相比监督学习基线提升2.2%至15.7%。结果为未来可落地的无线感知系统设计提供了重要方向。

原文摘要 · Abstract (English)

While Wi-Fi sensing offers a compelling, privacy-preserving alternative to cameras, its practical utility has been fundamentally undermined by a lack of robustness across domains. Models trained in one setup fail to generalize to new environments, hardware, or users, a critical "domain shift" problem exacerbated by modest, fragmented public datasets. We shift from this limited paradigm and apply a foundation model approach, leveraging Masked Autoencoding (MAE) style pretraining on the largest and most heterogeneous Wi-Fi CSI datasets collection assembled to date. Our study pretrains and evaluates models on over 1.3 million samples extracted from 14 datasets, collected using 4 distinct devices across the 2.4/5/6 GHz bands and bandwidths from 20 to 160 MHz. Our large-scale evaluation is the first to systematically disentangle the impacts of data diversity versus model capacity on cross-domain performance. The results establish scaling trends on Wi-Fi CSI sensing. First, our experiments show log-linear improvements in unseen domain performance as the amount of pretraining data increases, suggesting that data scale and diversity are key to domain generalization. Second, based on the current data volume, larger model can only provide marginal gains for cross-domain performance, indicating that data, rather than model capacity, is the current bottleneck for Wi-Fi sensing generalization. Finally, we conduct a series of cross-domain evaluations on human activity recognition, human gesture recognition and user identification tasks. The results show that the large-scale pretraining improves cross-domain accuracy ranging from 2.2% to 15.7%, compared to the supervised learning baseline. Overall, our findings provide insightful direction for designing future Wi-Fi sensing systems that can eventually be robust enough for real-world deployment.

Wi-Fi感知跨域泛化基础模型数据规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。