用文字提问引导图像信息提取,提升复杂场景通信效率
Multi-Modal Semantic Communication
- 引入文本查询引导视觉注意力,实现跨模态信息聚焦
- 根据信道带宽自适应传输图像块,总码率匹配信道容量
- 适合远程交互、增强现实等需精准传输的低带宽场景
语义通信旨在传输对任务最相关的信息而非原始数据,可显著提升远距呈现、增强现实和遥感等应用的通信效率。现有基于Transformer的方法利用自注意力图识别图像中的重要区域,但在多物体复杂场景中因缺乏显式任务引导而表现不佳。为此,本文提出一种新型多模态语义通信框架,通过文本用户查询引导信息提取过程。系统采用跨模态注意力机制融合视觉特征与语言嵌入,生成视觉数据的软相关性评分。基于这些评分及瞬时信道带宽,使用算法以自适应分辨率传输图像块,采用独立训练的编码器-解码器对,总码率匹配信道容量。接收端将图像块重建并拼合,以保留任务关键信息。该灵活且目标驱动的设计使复杂与带宽受限环境下的高效语义通信成为可能。
原文摘要 · Abstract (English)
Semantic communication aims to transmit information most relevant to a task rather than raw data, offering significant gains in communication efficiency for applications such as telepresence, augmented reality, and remote sensing. Recent transformer-based approaches have used self-attention maps to identify informative regions within images, but they often struggle in complex scenes with multiple objects, where self-attention lacks explicit task guidance. To address this, we propose a novel Multi-Modal Semantic Communication framework that integrates text-based user queries to guide the information extraction process. Our proposed system employs a cross-modal attention mechanism that fuses visual features with language embeddings to produce soft relevance scores over the visual data. Based on these scores and the instantaneous channel bandwidth, we use an algorithm to transmit image patches at adaptive resolutions using independently trained encoder-decoder pairs, with total bitrate matching the channel capacity. At the receiver, the patches are reconstructed and combined to preserve task-critical information. This flexible and goal-driven design enables efficient semantic communication in complex and bandwidth-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。