用图文多模态信息指导图像生成,提升语义通信的还原质量。
Image Generation with Supervised Selection Based on Multimodal Features for Semantic Communications
- 融合图像与文本特征,动态选择最优生成图像。
- 在低信噪比下仍保持高保真度,重建误差更低。
- 适合需要精准视觉还原的多用户通信场景。
语义通信(SemCom)是下一代通信系统的重要方向,允许接收端基于语义特征恢复源信号。然而,现有研究多仅依赖单一类型语义信息(如文本或图像)进行生成监督,难以全面捕捉准确语义,导致性能瓶颈。为此,本文提出一种基于多模态特征的语义通信框架,利用卷积神经网络(CNN)和对比语言-图像预训练模型(CLIP)分别提取图像与文本的语义特征。接收端采用生成式扩散模型生成多幅图像,并综合图像与文本特征,以最小重建误差选择最优图像。进一步将该多模态语义通信(MMSemCom)系统扩展至多用户正交传输场景。实验表明,所提框架在图像传输中显著提升保真度与鲁棒性,尤其在低信噪比(SNR)条件下表现优异。
原文摘要 · Abstract (English)
Semantic communication (SemCom) has emerged as a promising technique for the next-generation communication systems, in which the generation at the receiver side is allowed with semantic features' recovery. However, the majority of existing research predominantly utilizes a singular type of semantic information, such as text, images, or speech, to supervise and choose the generated source signals, which may not sufficiently encapsulate the comprehensive and accurate semantic information, and thus creating a performance bottleneck. In order to bridge this gap, in this paper, we propose and investigate a SemCom framework using multimodal information to supervise the generated image. To be specific, in this framework, we first extract semantic features at both the image and text levels utilizing the Convolutional Neural Network (CNN) architecture and the Contrastive Language-Image Pre-Training (CLIP) model before transmission. Then, we employ a generative diffusion model at the receiver to generate multiple images. In order to ensure the accurate extraction and facilitate high-fidelity image reconstruction, we select the "best" image with the minimum reconstruction errors by taking both the aided image and text semantic features into account. We further extend multimodal semantic communication (MMSemCom) system to the multiuser scenario for orthogonal transmission. Experimental results demonstrate that the proposed framework can not only achieve the enhanced fidelity and robustness in image transmission compared with existing communication systems but also sustain a high performance in the low signal-to-noise ratio (SNR) conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。