arXiv:2505.10003cs.LGeess.SP2025-05被引 10

用大模型统一处理无线多模态任务,6G系统新范式

AI2MMUM: AI-AI Oriented Multi-Modal Universal Model Leveraging Telecom Domain Large Model

  • 以大模型为基底,融合语言与无线模态,支持灵活任务指令
  • 在WAIR-D和DeepMIMO数据集上五项任务达最新最佳性能
  • 适合研究6G物理层智能与多模态系统设计的学者

面向6G的通用模型需具备处理多模态数据并执行多样空中接口任务的能力。基于先前的通信多模态对齐与电信大语言模型工作,我们提出可扩展、任务感知的人工智能-空气接口多模态通用模型(AI2MMUM),能根据细微任务指令灵活高效地完成各类物理层任务。以大语言模型为骨干,提供强上下文理解与泛化能力,并通过微调融入领域知识。任务指令包含固定关键词与可学习隐式前缀提示,提升适应性。冻结的射频模态编码器提取通用表征,适配层连接射频与语言模态,轻量级任务头直接输出目标结果。全面评估表明,AI2MMUM在使用WAIR-D与DeepMIMO数据集的五项代表性物理环境/无线信道下游任务中达到当前最优性能。

原文摘要 · Abstract (English)

Designing a 6G-oriented universal model capable of processing multi-modal data and executing diverse air interface tasks has emerged as a common goal in future wireless systems. Building on our prior work in communication multi-modal alignment and telecom large language model (LLM), we propose a scalable, task-aware artificial intelligence-air interface multi-modal universal model (AI2MMUM), which flexibility and effectively perform various physical layer tasks according to subtle task instructions. The LLM backbone provides robust contextual comprehension and generalization capabilities, while a fine-tuning approach is adopted to incorporate domain-specific knowledge. To enhance task adaptability, task instructions consist of fixed task keywords and learnable, implicit prefix prompts. Frozen radio modality encoders extract universal representations and adapter layers subsequently bridge radio and language modalities. Moreover, lightweight task-specific heads are designed to directly output task objectives. Comprehensive evaluations demonstrate that AI2MMUM achieves SOTA performance across five representative physical environment/wireless channel-based downstream tasks using the WAIR-D and DeepMIMO datasets.

6G多模态大模型无线系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。