arXiv:2608.14260eess.IVcs.MM2026-08中稿 · IEEE GLOBECOM 2026

用视觉语言模型实现图像传输的个性化语义通信,更贴合用户偏好。

Personalized Digital Semantic Communication for Image Transmission with Vision-Language Models

论文配图:Personalized Digital Semantic Communication for Image Transmission with Vision-Language Models
图 1 · 摘自论文原文
  • 基于视觉语言模型提取源图像与用户历史交互的个性化语义特征
  • 在有限带宽下比现有方法更保真且更符合用户偏好
  • 适合对用户体验敏感的无线图像传输场景

语义通信(SC)可实现高效无线图像传输,但现有方案多为无用户差异设计,忽略接收端语义需求。为此,我们提出个性化数字语义通信(PDSC)框架,结合基于视觉语言模型(VLM)的语义编码器与基于潜在扩散模型(LDM)的语义解码器。编码器从源图像及接收方历史交互中提取源感知的个性化语义标记,经向量量化为离散索引并编码为固定长度比特流,兼容数字传输。解码器根据恢复的语义标记重建个性化图像。此外,我们构建了容量受限的个性化语义率失真问题,并引入联合刻画源语义保真度与用户偏好一致性的语义失真度量。实验表明,在带宽受限无线传输下,PDSC 在源语义一致性与个性化方面优于当前先进基线方法,包括 CDDM 与 MoS。

原文摘要 · Abstract (English)

Semantic communication (SC) enables bandwidth-efficient wireless image transmission, but most existing SC schemes are user-agnostic and ignore receiver-dependent semantics. To address this issue, we propose a personalized digital semantic communication (PDSC) framework that integrates a vision-language model (VLM)-based semantic encoder with a latent diffusion model (LDM)-based semantic decoder. Specifically, the semantic encoder extracts source-aware personalized semantic tokens from both the source image and the receiver's historical interactions. These tokens are vector-quantized into discrete semantic indices and further encoded into a compact fixed-length bitstream, enabling compatibility with digital transmission. At the receiver, the semantic decoder reconstructs a personalized image conditioned on the recovered semantic tokens. Furthermore, we formulate a capacity-constrained personalized semantic rate-distortion problem and introduce a semantic distortion metric that jointly characterizes source-semantic fidelity and user-preference alignment. Experiments show that PDSC achieves superior source-semantic consistency and personalization over state-of-the-art SC baselines, including CDDM and MoS, under bandwidth-limited wireless transmission.

语义通信视觉语言模型个性化传输图像压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。