arXiv:2412.06038eess.SPcs.CV2024-12被引 25

用视觉变压器注意力量化图像块重要性,实现无需端到端训练的高效无线图像语义传输。

Vision Transformer-based Semantic Communications With Importance-Aware Quantization

  • 基于预训练ViT的注意力得分动态分配图像块量化比特数。
  • 在无错误和真实信道下均优于传统图像压缩方法,多视图分类准确率提升显著。
  • 适合无线图像传输场景,尤其适用于资源受限的智能设备部署。

语义通信通过传输与任务相关的语义特征,在无线信道中实现显著性能提升。然而,现有研究大多依赖神经编码器与解码器的端到端训练以保证语义特征的有效传输。为摆脱对端到端训练的依赖,本文提出一种基于视觉变换器(ViT)的语义通信系统,结合重要性感知量化(IAQ)用于无线图像传输。核心思想是利用预训练ViT模型的注意力得分来量化图像块的重要性,并据此为不同块分配不同的量化比特。为此,我们构建了一个加权量化误差最小化问题,权重设为注意力得分的递增函数,进而设计了最优增量分配法与低复杂度水填法求解。该框架进一步通过等效二进制对称信道(BSC)模型扩展至实际数字通信系统。在单视图与多视图图像分类任务上的仿真表明,本方法在无错与真实通信场景下均优于传统图像压缩方法。

原文摘要 · Abstract (English)

Semantic communications provide significant performance gains over traditional communications by transmitting task-relevant semantic features through wireless channels. However, most existing studies rely on end-to-end (E2E) training of neural-type encoders and decoders to ensure effective transmission of these semantic features. To enable semantic communications without relying on E2E training, this paper presents a vision transformer (ViT)-based semantic communication system with importance-aware quantization (IAQ) for wireless image transmission. The core idea of the presented system is to leverage the attention scores of a pretrained ViT model to quantify the importance levels of image patches. Based on this idea, our IAQ framework assigns different quantization bits to image patches based on their importance levels. This is achieved by formulating a weighted quantization error minimization problem, where the weight is set to be an increasing function of the attention score. Then, an optimal incremental allocation method and a low-complexity water-filling method are devised to solve the formulated problem. Our framework is further extended for realistic digital communication systems by modifying the bit allocation problem and the corresponding allocation methods based on an equivalent binary symmetric channel (BSC) model. Simulations on single-view and multi-view image classification tasks show that our IAQ framework outperforms conventional image compression methods in both error-free and realistic communication scenarios.

语义通信视觉变换器量化无线传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。