arXiv:2410.19766cs.CVcs.LG2024-10被引 26

用视觉大模型知识提升小样本射频人体活动识别效果

Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition

  • 跨模态对比知识蒸馏,让射频模型学会视觉大模型的语义理解
  • 零样本学习下性能媲美视觉方法,支持多种环境泛化
  • 适合数据稀缺的射频感知场景,如隐私敏感应用

基于射频(RF)的人体活动识别(HAR)为无法使用计算机视觉的应用提供了有前景的解决方案。然而,由于标签数据稀少且难以解释,限制了其发展。得益于基础模型(FMs)的突破,从无标签视觉数据中提取深层语义成为可能,但这些视觉基础模型在小规模射频数据集上表现不佳。为此,我们提出FM-Fi,一种创新的跨模态框架,将视觉基础模型的知识迁移至射频领域以增强HAR系统。FM-Fi采用新颖的跨模态对比知识蒸馏机制,使射频编码器能够继承基础模型的解释能力,实现零样本学习;同时利用基础模型与射频数据的内在特性,去除冗余特征,促进模态间更好对齐。框架进一步通过基于度量的少样本学习技术优化,旨在提升预定义HAR任务的性能。全面评估表明,FM-Fi在效果上可媲美视觉方法,验证了其在多种环境下的泛化能力。

原文摘要 · Abstract (English)

Radio-Frequency (RF)-based Human Activity Recognition (HAR) rises as a promising solution for applications unamenable to techniques requiring computer visions. However, the scarcity of labeled RF data due to their non-interpretable nature poses a significant obstacle. Thanks to the recent breakthrough of foundation models (FMs), extracting deep semantic insights from unlabeled visual data become viable, yet these vision-based FMs fall short when applied to small RF datasets. To bridge this gap, we introduce FM-Fi, an innovative cross-modal framework engineered to translate the knowledge of vision-based FMs for enhancing RF-based HAR systems. FM-Fi involves a novel cross-modal contrastive knowledge distillation mechanism, enabling an RF encoder to inherit the interpretative power of FMs for achieving zero-shot learning. It also employs the intrinsic capabilities of FM and RF to remove extraneous features for better alignment between the two modalities. The framework is further refined through metric-based few-shot learning techniques, aiming to boost the performance for predefined HAR tasks. Comprehensive evaluations evidently indicate that FM-Fi rivals the effectiveness of vision-based methodologies, and the evaluation results provide empirical validation of FM-Fi's generalizability across various environments.

射频识别跨模态小样本学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。