arXiv:2505.18194eess.SPcs.AI2025-05被引 5

用大模型协同多设备感知,提升复杂场景下的精度与隐私保护。

Large Language Model-Driven Distributed Integrated Multimodal Sensing and Semantic Communications

  • 多设备融合射频与视觉数据,通过交叉注意力实现跨模态整合。
  • 在合成数据集上,系统感知准确率显著优于单模态方法。
  • 支持分布式训练,兼顾隐私保护与高效语义通信,适合智慧城市应用。

传统单模态感知系统(仅依赖射频或视觉数据)难以应对复杂动态环境,且单设备视角有限、覆盖不足,影响城市或非视距场景下的性能。为此,提出一种大语言模型驱动的分布式集成多模态感知与语义通信框架(LLM-DiSAC)。系统由多个配备射频与摄像头的协同感知设备及一个汇聚中心组成。首先,在感知设备端,设计射频-视觉融合网络(RVFN),分别使用专用特征提取器处理射频与视觉数据,并通过交叉注意力模块实现有效融合。其次,提出基于大语言模型的语义传输网络(LSTN),利用已知信道参数(如收发距离、信噪比)减少语义失真。第三,在汇聚中心,采用基于Transformer的聚合模型(TRAM),结合自适应聚合注意力机制融合分布式特征,提升感知精度。为保障数据隐私,引入两阶段分布式学习策略:设备端本地训练,汇聚中心基于中间特征进行集中式模型训练。在由Genesis仿真引擎生成的合成多视角射频-视觉数据集上评估表明,该系统表现优异。

原文摘要 · Abstract (English)

Traditional single-modal sensing systems-based solely on either radio frequency (RF) or visual data-struggle to cope with the demands of complex and dynamic environments. Furthermore, single-device systems are constrained by limited perspectives and insufficient spatial coverage, which impairs their effectiveness in urban or non-line-of-sight scenarios. To overcome these challenges, we propose a novel large language model (LLM)-driven distributed integrated multimodal sensing and semantic communication (LLM-DiSAC) framework. Specifically, our system consists of multiple collaborative sensing devices equipped with RF and camera modules, working together with an aggregation center to enhance sensing accuracy. First, on sensing devices, LLM-DiSAC develops an RF-vision fusion network (RVFN), which employs specialized feature extractors for RF and visual data, followed by a cross-attention module for effective multimodal integration. Second, a LLM-based semantic transmission network (LSTN) is proposed to enhance communication efficiency, where the LLM-based decoder leverages known channel parameters, such as transceiver distance and signal-to-noise ratio (SNR), to mitigate semantic distortion. Third, at the aggregation center, a transformer-based aggregation model (TRAM) with an adaptive aggregation attention mechanism is developed to fuse distributed features and enhance sensing accuracy. To preserve data privacy, a two-stage distributed learning strategy is introduced, allowing local model training at the device level and centralized aggregation model training using intermediate features. Finally, evaluations on a synthetic multi-view RF-visual dataset generated by the Genesis simulation engine show that LLM-DiSAC achieves a good performance.

多模态感知大模型分布式系统语义通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。