用大模型分区域传图,关键部分更清晰。
Large AI Model-Enabled Generative Semantic Communications for Image Transmission
- 图像分关键区与非关键区,分别用不同方式传输
- 相比传统方法,视觉质量和语义保真度均提升
- 轻量化部署让大模型也能高效运行
生成式人工智能的快速发展为语义通信系统中的图像传输效率与精度带来了新机遇。然而,现有方法常忽视图像各区域的重要性差异,可能影响视觉关键内容的重建质量。为此,本文提出一种创新的生成式语义通信系统,通过将图像分割为关键与非关键区域,实现更精细的语义粒度控制:关键区域采用面向图像的语义编码器处理,非关键区域则通过图像到文本建模方法高效压缩。此外,为缓解大型AI模型带来的存储与计算压力,系统采用轻量化部署策略,结合模型量化与低秩自适应微调技术,在不损失性能的前提下显著提升资源利用率。仿真结果表明,该系统在语义保真度和视觉质量方面均优于传统方法,验证了其在图像传输任务中的有效性。
原文摘要 · Abstract (English)
The rapid development of generative artificial intelligence (AI) has introduced significant opportunities for enhancing the efficiency and accuracy of image transmission within semantic communication systems. Despite these advancements, existing methodologies often neglect the difference in importance of different regions of the image, potentially compromising the reconstruction quality of visually critical content. To address this issue, we introduce an innovative generative semantic communication system that refines semantic granularity by segmenting images into key and non-key regions. Key regions, which contain essential visual information, are processed using an image oriented semantic encoder, while non-key regions are efficiently compressed through an image-to-text modeling approach. Additionally, to mitigate the substantial storage and computational demands posed by large AI models, the proposed system employs a lightweight deployment strategy incorporating model quantization and low-rank adaptation fine-tuning techniques, significantly boosting resource utilization without sacrificing performance. Simulation results demonstrate that the proposed system outperforms traditional methods in terms of both semantic fidelity and visual quality, thereby affirming its effectiveness for image transmission tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。