arXiv:2510.13390cs.CV2025-10中稿 · IEEE ICPADS 2025被引 2

用大模型提升WiFi手势识别的泛化能力,更准更快更省资源。

Generalizing WiFi Gesture Recognition via Large-Model-Aware Semantic Distillation and Alignment

  • 利用大模型语义先验,通过双路径编码提取手势几何与动态特征。
  • 在Widar3.0上跨域识别准确率超现有方法,模型体积减小57%。
  • 适合部署在资源受限的AIoT设备,如智能家居、可穿戴终端。

基于WiFi的手势识别已成为实现AIoT环境中非接触、隐私保护人机交互的有前景的射频感知范式。然而,现有方法常因信道状态信息(CSI)的领域敏感性及缺乏高层语义抽象而面临泛化能力弱、语义表达不足的问题。为此,本文提出一种新型泛化框架GLSDA,利用预训练大模型的语义先验,在域内与跨域场景中增强手势表征学习。首先设计双路径CSI编码流程,通过CSI比值相位序列和多普勒谱图捕捉几何与动态手势模式;再通过多尺度语义编码器学习鲁棒时序嵌入,并借助跨模态注意力对齐语义。为增强类别判别力,引入语义感知软监督机制,编码类间相关性,缓解语义相似手势的标签歧义。最后,采用鲁棒双蒸馏策略,将对齐模型压缩为轻量学生网络,联合蒸馏中间特征与语义引导的软标签。在Widar3.0基准上的大量实验表明,GLSDA在域内与跨域任务中均持续优于当前最优方法,同时显著降低模型大小与推理延迟。本方法为真实世界AIoT应用中的通用射频手势接口提供了可扩展、可部署的解决方案。

原文摘要 · Abstract (English)

WiFi-based gesture recognition has emerged as a promising RF sensing paradigm for enabling non-contact and privacy-preserving human-computer interaction in AIoT environments. However, existing methods often suffer from limited generalization and semantic expressiveness due to the domain-sensitive nature of Channel State Information and the lack of high-level gesture abstraction. To address these challenges, we propose a novel generalization framework, termed Large-Model-Aware Semantic Distillation and Alignment (GLSDA), which leverages the semantic prior of pre-trained large foundation models to enhance gesture representation learning in both in-domain and cross-domain scenarios. Specifically, we first design a dual-path CSI encoding pipeline that captures geometric and dynamic gesture patterns via CSI-Ratio phase sequences and Doppler spectrograms. These representations are then fed into a Multiscale Semantic Encoder, which learns robust temporal embeddings and aligns them with gesture semantics through cross-modal attention mechanisms. To further enhance category discrimination, we introduce a Semantic-Aware Soft Supervision scheme that encodes inter-class correlations and reduces label ambiguity, especially for semantically similar gestures. Finally, we develop a Robust Dual-Distillation strategy to compress the aligned model into a lightweight student network, jointly distilling intermediate features and semantic-informed soft labels from the teacher model. Extensive experiments on the Widar3.0 benchmark show that GLSDA consistently outperforms state-of-the-art methods in both in-domain and cross-domain gesture recognition tasks, while significantly reducing model size and inference latency. Our method offers a scalable and deployable solution for generalized RF-based gesture interfaces in real-world AIoT applications.

WiFi感知手势识别大模型蒸馏AIoT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。