arXiv:2608.00642cs.CV2026-08

融合信道幅值与多普勒-时延特征,提升无线信号下人体动作识别精度

WiFuse: An Attention Mechanism for Human Activity Recognition using Fused CSI Amplitude and Delay-Doppler Channel Features

论文配图:WiFuse: An Attention Mechanism for Human Activity Recognition using Fused CSI Amplitude and Delay-Doppler Channel Features
图 1 · 摘自论文原文
  • 双流架构分别处理时域幅值和相位导出的多普勒-时延图
  • 在XRF55和Wi-MIR数据集上准确率分别达95.28%和98.20%
  • 适用于复杂环境下的隐私保护动作识别,尤其抗多径与噪声干扰

近年来,无线感知在人体活动识别(HAR)中扮演重要角色,仅通过无线信号即可检测多种动作,保障用户隐私且无侵入性。然而,反射面、硬件偏移及其他物理损伤会干扰神经网络,导致误判并显著降低模型准确率。为此,我们提出WiFuse框架,一种基于双流信道状态信息(CSI)的人体活动识别方法,将去噪后的时域幅值变化与经净化相位计算得到的二维快速傅里叶变换(2D-FFT)延迟-多普勒运动表示进行融合。该融合表征输入混合残差网络(ResNet)与时间卷积网络(TCN)结构,其中ResNet提取空间-频谱特征,TCN建模长程时间依赖,并引入通道与时空注意力机制;采用解耦的两阶段迁移学习策略以提升优化稳定性与特征复用效率。我们在两个公开数据集上进行了广泛实验,包括与最先进方法及其它混合架构的对比、消融研究、跨数据集测试与域适应评估。所提框架在XRF55数据集四个环境下的整体准确率最高达95.28%,在多用户Wi-MIR数据集上达到98.20%。结果表明,在通常导致深度神经网络性能下降的条件下(如类别重叠、多径传播、噪声与干扰),结合幅值与延迟-多普勒表示的双流策略,辅以迁移学习,能有效提升识别性能。

原文摘要 · Abstract (English)

Recently, Wi-Fi sensing has played a significant role in Human Activity Recognition (HAR), as it enables the detection of various activities using only Wi-Fi signals, ensuring privacy and remaining non-intrusive for the user. However, environmental characteristics such as reflective surfaces, hardware offsets, and other physical impairments affect recognition by the neural network, subsequently causing errors and significantly reducing model accuracy. To overcome this problem we present the WiFuse framework, a dual-stream Channel State Information (CSI) framework for human activity recognition (HAR) that pairs denoised time-domain amplitude variations with 2D-FFT-derived Delay-Doppler motion representations computed from the sanitized channel phase. The fused representation feeds a hybrid ResNet-Temporal Convolutional Network (TCN) neural architecture augmented with channel and spatio-temporal attention, where the ResNet extracts spatial-spectral features and the TCN models long-range temporal dependencies; a decoupled two-stage transfer learning strategy is employed to improve optimization stability and feature reuse. We conduct extensive experiments on two public datasets, including comparisons against state-of-the-art methods and alternative hybrid architectures, ablation studies, and cross-dataset and domain-adaptation evaluations. The proposed framework reaches an overall accuracy of up to 95.28% across the four environments of the XRF55 dataset and up to 98.20% on the multi-user Wi-MIR dataset. Overall, the results indicate that combining amplitude and Delay-Doppler representations within a dual-stream strategy, enhanced by transfer learning, improves recognition performance under conditions that typically degrade deep neural networks, such as class overlap, multipath propagation, noise, and interference.

人体动作识别无线感知双流网络注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。