arXiv:2602.14833eess.SPcs.LG2026-02被引 11

让AI理解无线信号,用视觉模型分析频谱图实现多任务识别

RF-GPT: Teaching AI to See the Wireless World

  • 用多模态模型的视觉编码器处理无线频谱图,生成射频标记输入语言模型
  • 在6种无线技术上构建1.2万场景、62.5万指令数据,零人工标注训练
  • 可同时完成调制识别、信号重叠分析、5G信息提取等任务,通用模型无法胜任

大语言模型(LLM)和多模态模型已成为强大的通用推理系统,但支持无线通信的射频(RF)信号仍无法被这些模型原生处理。现有基于LLM的电信方法主要依赖文本与结构化数据,而传统深度学习的射频模型则针对特定信号处理任务独立构建,暴露出射频感知与高层推理之间的明显鸿沟。为此,我们提出射频语言模型(RFLM)RF-GPT,利用多模态大模型的视觉编码器处理射频频谱图。复杂正交波形(IQ)被映射为时频谱图,输入预训练视觉编码器,得到的特征作为射频标记注入仅解码器的LLM,生成射频相关的回答、解释与结构化输出。为训练RF-GPT,我们使用完全合成的射频语料对预训练多模态模型进行监督指令微调。基于符合标准的波形生成器,构建了涵盖六种无线技术的宽带场景,从中提取时频谱图、精确配置元数据和密集描述。文本模型将描述转化为射频接地的指令-答案对,共生成约1.2万射频场景与62.5万条指令样本,无需人工标注。在宽带调制分类、信号重叠分析、无线技术识别、WLAN用户计数及5G NR信息提取等多个基准测试中,RF-GPT展现出强大多任务性能,而无射频接地的通用视觉语言模型则表现不佳。

原文摘要 · Abstract (English)

Large language models (LLMs) and multimodal models have become powerful general-purpose reasoning systems. However, radio-frequency (RF) signals, which underpin wireless systems, are still not natively supported by these models. Existing LLM-based approaches for telecom focus mainly on text and structured data, while conventional RF deep-learning models are built separately for specific signal-processing tasks, highlighting a clear gap between RF perception and high-level reasoning. To bridge this gap, we introduce RF-GPT, a radio-frequency language model (RFLM) that utilizes the visual encoders of multimodal LLMs to process and understand RF spectrograms. In this framework, complex in-phase/quadrature (IQ) waveforms are mapped to time-frequency spectrograms and then passed to pretrained visual encoders. The resulting representations are injected as RF tokens into a decoder-only LLM, which generates RF-grounded answers, explanations, and structured outputs. To train RF-GPT, we perform supervised instruction fine-tuning of a pretrained multimodal LLM using a fully synthetic RF corpus. Standards-compliant waveform generators produce wideband scenes for six wireless technologies, from which we derive time-frequency spectrograms, exact configuration metadata, and dense captions. A text-only LLM then converts these captions into RF-grounded instruction-answer pairs, yielding roughly 12,000 RF scenes and 0.625 million instruction examples without any manual labeling. Across benchmarks for wideband modulation classification, overlap analysis, wireless-technology recognition, WLAN user counting, and 5G NR information extraction, RF-GPT achieves strong multi-task performance, whereas general-purpose VLMs with no RF grounding largely fail.

射频感知多模态模型语言模型无线通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。