用视觉Transformer实现高效语义通信,图像传输更省带宽且保真度高。
Vision Transformer Based Semantic Communications for Next Generation Wireless Networks
- 采用视觉Transformer构建编码解码框架,专注提取图像语义信息。
- 在真实信道环境下仍保持38 dB的峰值信噪比,优于传统深度学习方法。
- 适合6G时代低延迟、高效率的语义通信场景,尤其关注带宽敏感应用。
在6G网络演进背景下,语义通信有望通过优先传输语义信息而非原始数据来革新数据传输方式。本文提出一种基于视觉Transformer(ViT)的语义通信框架,旨在实现图像传输中高语义相似性的同时显著降低带宽需求。通过将ViT作为编码器-解码器架构,该方案可在发射端高效编码图像以保留高语义内容,并在接收端考虑实际信道衰落与噪声的情况下精确重建图像。得益于ViT固有的注意力机制,该模型在性能上超越了针对此类图像生成任务优化的卷积神经网络(CNN)和生成对抗网络(GAN)。所提出的基于ViT的架构在不同通信环境中实现了38 dB的峰值信噪比(PSNR),高于其他深度学习方法,充分证明其在保持语义一致性方面的优越性,为语义通信领域带来重要突破。
原文摘要 · Abstract (English)
In the evolving landscape of 6G networks, semantic communications are poised to revolutionize data transmission by prioritizing the transmission of semantic meaning over raw data accuracy. This paper presents a Vision Transformer (ViT)-based semantic communication framework that has been deliberately designed to achieve high semantic similarity during image transmission while simultaneously minimizing the demand for bandwidth. By equipping ViT as the encoder-decoder framework, the proposed architecture can proficiently encode images into a high semantic content at the transmitter and precisely reconstruct the images, considering real-world fading and noise consideration at the receiver. Building on the attention mechanisms inherent to ViTs, our model outperforms Convolution Neural Network (CNNs) and Generative Adversarial Networks (GANs) tailored for generating such images. The architecture based on the proposed ViT network achieves the Peak Signal-to-noise Ratio (PSNR) of 38 dB, which is higher than other Deep Learning (DL) approaches in maintaining semantic similarity across different communication environments. These findings establish our ViT-based approach as a significant breakthrough in semantic communications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。