融合雷达与图像的语义感知通信框架,实现多任务高效协同
SIMAC: A Semantic-Driven Integrated Multimodal Sensing And Communication Framework
- 通过跨模态注意力融合雷达与图像语义特征
- 支持多任务感知,实验显示精度更高
- 适合需要多场景感知的智能系统
传统单模态感知在精度和能力上受限,且与通信系统分离导致带宽受限环境下延迟增加。同时,单任务感知难以满足用户多样化需求。为此,我们提出语义驱动的集成多模态感知与通信框架(SIMAC)。该框架采用联合信源信道编码架构,实现感知结果的同步解码与传输。首先引入多模态语义融合(MSF)网络,分别提取雷达信号与图像的语义信息,并通过交叉注意力机制融合生成多模态语义表示。其次提出基于大语言模型(LLM)的语义编码器(LSE),将通信参数与多模态语义映射至统一潜在空间并输入LLM,实现信道自适应语义编码。第三,设计面向任务的感知语义解码器(SSD),根据不同任务需求配置解码头。同时引入多任务学习策略训练整个框架,实现多样化的感知服务。最终实验仿真表明,该框架可支持多种感知任务,且具备更高精度。
原文摘要 · Abstract (English)
Traditional single-modality sensing faces limitations in accuracy and capability, and its decoupled implementation with communication systems increases latency in bandwidth-constrained environments. Additionally, single-task-oriented sensing systems fail to address users' diverse demands. To overcome these challenges, we propose a semantic-driven integrated multimodal sensing and communication (SIMAC) framework. This framework leverages a joint source-channel coding architecture to achieve simultaneous sensing decoding and transmission of sensing results. Specifically, SIMAC first introduces a multimodal semantic fusion (MSF) network, which employs two extractors to extract semantic information from radar signals and images, respectively. MSF then applies cross-attention mechanisms to fuse these unimodal features and generate multimodal semantic representations. Secondly, we present a large language model (LLM)-based semantic encoder (LSE), where relevant communication parameters and multimodal semantics are mapped into a unified latent space and input to the LLM, enabling channel-adaptive semantic encoding. Thirdly, a task-oriented sensing semantic decoder (SSD) is proposed, in which different decoded heads are designed according to the specific needs of tasks. Simultaneously, a multi-task learning strategy is introduced to train the SIMAC framework, achieving diverse sensing services. Finally, experimental simulations demonstrate that the proposed framework achieves diverse sensing services and higher accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。