用统一视觉语言特征实现图文同步传输,省带宽还抗干扰。
VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System
- 用预训练多模态模型生成紧凑的视觉语言特征,统一编码图文信息。
- 在低信噪比下仍保持高语义准确率,比纯文本/图像传输节省大量带宽。
- 适合需要高效跨模态通信的场景,如远程协作、智能设备交互。
我们提出视觉语言特征为基础的多模态语义通信系统(VLF-MSC),通过传输单一紧凑的视觉语言表征,在接收端同时支持图像与文本生成。不同于传统分模态处理方式,VLF-MSC利用预训练视觉语言模型(VLM)将源图像编码为视觉语言语义特征(VLF),经无线信道传输。接收端使用基于解码器的语言模型和基于扩散的图像生成器,均以该VLF为条件,生成描述性文本与语义对齐图像。该统一表示避免了模态专用流或重传,提升频谱效率与适应性。借助基础模型,系统在信道噪声下仍保持语义保真度。实验表明,相较于仅文本或仅图像的基线,VLF-MSC在低信噪比条件下实现了更高的双模态语义准确率,并显著降低带宽消耗。
原文摘要 · Abstract (English)
We propose Vision-Language Feature-based Multimodal Semantic Communication (VLF-MSC), a unified system that transmits a single compact vision-language representation to support both image and text generation at the receiver. Unlike existing semantic communication techniques that process each modality separately, VLF-MSC employs a pre-trained vision-language model (VLM) to encode the source image into a vision-language semantic feature (VLF), which is transmitted over the wireless channel. At the receiver, a decoder-based language model and a diffusion-based image generator are both conditioned on the VLF to produce a descriptive text and a semantically aligned image. This unified representation eliminates the need for modality-specific streams or retransmissions, improving spectral efficiency and adaptability. By leveraging foundation models, the system achieves robustness to channel noise while preserving semantic fidelity. Experiments demonstrate that VLF-MSC outperforms text-only and image-only baselines, achieving higher semantic accuracy for both modalities under low SNR with significantly reduced bandwidth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。