用大模型提升6G沉浸式通信的语义传输效率
Multimodal LLM Integrated Semantic Communications for 6G Immersive Experiences
- 将多模态大模型融入通信框架,实现上下文感知的智能传输
- 在AR/VR场景下,显著提升语义信息优先级与内容重建质量
- 适合研究6G智能通信、多模态生成的工程师和学者
6G网络旨在实现增强现实(AR)、虚拟现实(VR)及全息通信等革命性沉浸式体验。这些应用需实时传输高维多模态数据并进行智能处理,对资源受限的无线系统构成巨大挑战。同时,理解环境、上下文与用户意图对交付任务相关的内容至关重要。本文提出一种新型多模态大语言模型(MLLM)集成语义通信框架——MLLM-SC,充分利用预训练基础模型的推理与生成能力,实现上下文感知、任务导向的无线通信。该框架采用设备-边缘协同架构:边缘侧的MLLM赋能语义引导模块,分析多模态输入、用户意图与信道状态,生成重要性感知的注意力图,以优先传输语义关键信息。设计了重要性感知语义编码器与资源自适应解码器,联合优化以实现动态带宽分配与高质量内容重建或生成。在AR/VR场景下的视觉问答与扩散驱动图像生成的大量案例研究验证了MLLM-SC的有效性。
原文摘要 · Abstract (English)
6G networks promise revolutionary immersive communication experiences including augmented reality (AR), virtual reality (VR), and holographic communications. These applications demand high-dimensional multimodal data transmission and intelligent data processing in real-time, which is extremely challenging over resource-limited wireless communication systems. Moreover, a joint understanding of the environment, context, and user intent is essential to deliver task-relevant content effectively. This article presents a novel multimodal large language model (MLLM) integrated semantic communications framework, termed MLLM-SC, which fully leverages reasoning and generative capabilities of pre-trained foundation models for context-aware and task-oriented wireless communication. The MLLM-SC framework adopts a device-edge collaborative architecture. At the edge, MLLM-empowered semantic guidance module analyzes multimodal inputs, user intents, and channel conditions to generate importance-aware attention maps prioritizing semantically critical information. An importance-aware semantic encoder and a resource-adaptive semantic decoder are jointly designed and optimized, which can utilize the semantic guidance for adaptive bandwidth allocation and high-quality content reconstruction or generation. Extensive case studies on visual question answering for AR/VR applications and diffusion-driven image generation validate the effectiveness of MLLM-SC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。