递归视觉变压器动态调节深度宽度,高效压缩图像语义通信
Recursive Vision Transformer with Dynamic Depth and Width Adjustment for Resource-Efficient Image Semantic Communication
- 用递归结构迭代优化语义特征,减少参数量
- 动态调整深度和宽度,降低48.7%参数量
- 适合资源受限设备部署,重建质量更优
图像语义通信是下一代无线通信系统的关键组件。然而,现有系统通常存在内存占用大、计算复杂度高的问题,难以在资源受限设备上部署。为此,我们提出一种基于视觉变压器(ViT)的图像语义通信系统。该系统引入递归结构,迭代细化语义特征并减少参数量。同时设计三种动态调整策略:动态深度调整、动态宽度调整及宽深联合优化。动态深度调整根据图像内容与信道条件自适应决定递归模块数量;动态宽度调整选择性保留重要神经元与注意力头;联合优化进一步实现灵活计算配置。仿真结果表明,所提出的递归ViT系统结合三项动态调整策略,在相近计算复杂度下,参数量减少48.7%,且重建质量优于现有基线。
原文摘要 · Abstract (English)
Image semantic communication is a critical component in next-generation wireless communication systems. However, such systems typically suffer from large memory footprints and high computational complexity, making them difficult to deploy on resource-constrained devices. To address these challenges, we propose a vision transformer (ViT)-enabled image semantic communication system. In this system, a recursive structure is introduced to iteratively refine semantic features and reduce the parameter count. In addition, three dynamic adjustment strategies are designed to adaptively reduce computational complexity: dynamic depth adjustment, dynamic width adjustment, and joint width-depth optimization. Dynamic depth adjustment adaptively determines the number of recursive modules according to image content and channel conditions, while dynamic width adjustment selectively preserves important neurons and attention heads. The joint width-depth optimization further enables flexible computation configurations. Simulation results verify that the proposed recursive ViT-based system, combined with the three dynamic adjustment strategies, reduces the parameter count by 48.7% and achieves higher reconstruction quality than existing baselines under comparable computational complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。