arXiv:2411.09996eess.SPcs.AI2024-11被引 19

用视觉变压器预训练频谱图,实现6G无线通信的通用模型。

Building 6G Radio Foundation Models with Transformer Architectures

  • 采用掩码频谱建模自监督预训练视觉变压器。
  • 在信道状态感知和频谱分割任务上性能媲美有监督训练。
  • 小模型胜过大模型,适合未来6G可扩展系统部署。

基础深度学习模型是通用模型,旨在学习目标模态的通用、鲁棒且可适应的表示,支持在多种下游任务中微调。这些模型通过自监督学习(SSL)在大规模无标签数据集上预训练。基础模型相比传统监督方法展现出更强的泛化能力,这对动态环境下的无线通信至关重要。本文提出并验证了视觉变压器(ViT)作为频谱图学习的无线电基础模型的有效性。我们引入掩码频谱建模(MSM)方法,以自监督方式预训练ViT。在两个下游任务上评估:基于信道状态信息(CSI)的人体活动感知与频谱图分割。实验结果表明,该模型性能可与有监督训练相媲美,并具备跨域泛化能力。值得注意的是,预训练的ViT模型在频谱分割任务上优于四倍大的从头训练模型,且训练时间显著更短;同时在基于CSI的人体活动感知任务上也达到竞争性表现。本工作证明了结合MSM的ViT在预训练方面的有效性,是未来6G网络中可扩展基础模型开发的有力候选技术。

原文摘要 · Abstract (English)

Foundation deep learning (DL) models are general models, designed to learn general, robust and adaptable representations of their target modality, enabling finetuning across a range of downstream tasks. These models are pretrained on large, unlabeled datasets using self-supervised learning (SSL). Foundation models have demonstrated better generalization than traditional supervised approaches, a critical requirement for wireless communications where the dynamic environment demands model adaptability. In this work, we propose and demonstrate the effectiveness of a Vision Transformer (ViT) as a radio foundation model for spectrogram learning. We introduce a Masked Spectrogram Modeling (MSM) approach to pretrain the ViT in a self-supervised fashion. We evaluate the ViT-based foundation model on two downstream tasks: Channel State Information (CSI)-based Human Activity sensing and Spectrogram Segmentation. Experimental results demonstrate competitive performance to supervised training while generalizing across diverse domains. Notably, the pretrained ViT model outperforms a four-times larger model that is trained from scratch on the spectrogram segmentation task, while requiring significantly less training time, and achieves competitive performance on the CSI-based human activity sensing task. This work demonstrates the effectiveness of ViT with MSM for pretraining as a promising technique for scalable foundation model development in future 6G networks.

6G视觉变压器自监督学习频谱建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。