arXiv:2503.19386cs.CVeess.SP2025-03

用视觉语言模型提取多文本语义,提升图像语义通信重建精度。

Exploring Textual Semantics Diversity for Image Transmission in Semantic Communication Systems using Visual Language Model

  • 将图像分块后用VLM生成多段文本描述,增强语义表达
  • 相比现有方法,重建准确率显著提升,验证了文本多样性有效性
  • 适合关注语义通信、多模态传输的科研与工程人员

近年来,机器学习的快速发展为传统通信系统带来革新与挑战。语义通信作为一种有效策略,能够提取图像相关的语义信号、分割标签和特征用于图像传输。然而,图像提取的语义特征数量不足可能导致重建精度偏低,制约其实际应用,仍是亟待解决的问题。为此,本文提出一种多文本传输语义通信(Multi-SC)系统,利用视觉语言模型(VLM)辅助图像语义信号传输。不同于以往系统,该方案将图像划分为多个块,通过改进的大型语言与视觉助手(LLaVA)提取多段文本信息,并结合语义分割标签与语义文本实现图像恢复。仿真结果表明,所提文本语义多样性方案显著提升了重建准确率,优于相关工作。

原文摘要 · Abstract (English)

In recent years, the rapid development of machine learning has brought reforms and challenges to traditional communication systems. Semantic communication has appeared as an effective strategy to effectively extract relevant semantic signals semantic segmentation labels and image features for image transmission. However, the insufficient number of extracted semantic features of images will potentially result in a low reconstruction accuracy, which hinders the practical applications and still remains challenging for solving. In order to fill this gap, this letter proposes a multi-text transmission semantic communication (Multi-SC) system, which uses the visual language model (VLM) to assist in the transmission of image semantic signals. Unlike previous image transmission semantic communication systems, the proposed system divides the image into multiple blocks and extracts multiple text information from the image using a modified large language and visual assistant (LLaVA), and combines semantic segmentation tags with semantic text for image recovery. Simulation results show that the proposed text semantics diversity scheme can significantly improve the reconstruction accuracy compared with related works.

语义通信视觉语言模型图像传输多文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。