基于大模型的图像语义编码,能按用户意图精准传输关注区域。
Generalized User-Oriented Image Semantic Coding Empowered by Large Vision-Language Model
- 根据用户文本查询提取图像相关特征,实现定向信息传输。
- 零样本测试下,问答匹配率优于现有最先进方法。
- 适合需要个性化、高泛化能力图像通信的场景。
语义通信在无线传输中表现出优异的整体信息保留能力。对于图像等语义丰富的内容,用户通常关注特定区域,取决于其意图。然而,现有语义编码模型多在特定数据集上训练,难以适应分布外的真实图像,泛化能力成为关键但未被充分探索的问题。为此,本文提出一种通用用户导向图像语义编码(UO-ISC)框架:用户输入文本查询以表达意图,发送端从源图像中提取与查询相关的特征,接收端基于这些特征重建图像。为增强泛化能力,引入预训练的大视觉语言模型(VLM)CLIP。为评估重建图像与用户查询的相关性,设计用户意图相关性损失,使用预训练的大语言-视觉助手模型LLaVA计算。在未见物体上的零样本推理实验表明,所提UO-ISC框架在问答匹配率上优于当前最先进的查询感知图像语义编码方法。
原文摘要 · Abstract (English)
Semantic communication has shown outstanding performance in preserving the overall source information in wireless transmission. For semantically rich content such as images, human users are often interested in specific regions depending on their intent. Moreover, recent semantic coding models are mostly trained on specific datasets. However, real-world applications may involve images out of the distribution of training dataset, which makes generalization a crucial but largely unexplored problem. To incorporate user's intent into semantic coding, in this paper, we propose a generalized user-oriented image semantic coding (UO-ISC) framework, where the user provides a text query indicating its intent. The transmitter extracts features from the source image which are relevant to the user's query. The receiver reconstructs an image based on those features. To enhance the generalization ability, we integrate contrastive language image pre-training (CLIP) model, which is a pretrained large vision-language model (VLM), into our proposed UO-ISC framework. To evaluate the relevance between the reconstructed image and the user's query, we introduce the user-intent relevance loss, which is computed by using a pretrained large VLM, large language-and-vision assistant (LLaVA) model. When performing zero-shot inference on unseen objects, simulation results show that the proposed UO-ISC framework outperforms the state-of-the-art query-aware image semantic coding in terms of the answer match rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。