用导频引导的多模态通信框架,提升音视频事件定位在复杂信道下的性能。
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
- 通过数字导频与动态信道状态结合,实现对真实信道的实时引导。
- 在变信道下仍保持优异性能,相比基准方法信噪比显著提升。
- 适用于自动驾驶、智能家居等需音视频协同定位的场景。
多模态语义通信融合文本、图像、音频等多种数据模态,显著提升通信效率与可靠性,在人工智能、自动驾驶和智能家居等领域具有广泛应用前景。然而,现有研究多依赖模拟信道并假设信道状态完美(理想信道状态信息),难以应对真实世界中动态变化的物理信道与噪声。当前方法通常仅关注单模态任务,无法有效处理视频与音频等多模态流数据及其对应任务。此外,现有的语义编码与解码模块主要传输单一模态特征,忽视了多模态语义增强与识别需求。为此,本文提出一种专为音视频事件定位设计的导频引导式多模态语义通信框架。该框架利用数字导频码与信道模块,引导真实场景中的模拟信道状态,并设计基于欧拉方法的多模态语义编码与解码机制,考虑时频特性并适应动态信道状态。该方法能有效处理多模态流数据,尤其在音视频事件定位任务中表现突出。大量数值实验表明,所提框架在信道变化下具有强鲁棒性,支持多种通信场景,且在信噪比(SNR)指标上优于现有基准方法,彰显其在语义通信质量上的优势。
原文摘要 · Abstract (English)
Multimodal semantic communication, which integrates various data modalities such as text, images, and audio, significantly enhances communication efficiency and reliability. It has broad application prospects in fields such as artificial intelligence, autonomous driving, and smart homes. However, current research primarily relies on analog channels and assumes constant channel states (perfect CSI), which is inadequate for addressing dynamic physical channels and noise in real-world scenarios. Existing methods often focus on single modality tasks and fail to handle multimodal stream data, such as video and audio, and their corresponding tasks. Furthermore, current semantic encoding and decoding modules mainly transmit single modality features, neglecting the need for multimodal semantic enhancement and recognition tasks. To address these challenges, this paper proposes a pilot-guided framework for multimodal semantic communication specifically tailored for audio-visual event localization tasks. This framework utilizes digital pilot codes and channel modules to guide the state of analog channels in real-wold scenarios and designs Euler-based multimodal semantic encoding and decoding that consider time-frequency characteristics based on dynamic channel state. This approach effectively handles multimodal stream source data, especially for audio-visual event localization tasks. Extensive numerical experiments demonstrate the robustness of the proposed framework in channel changes and its support for various communication scenarios. The experimental results show that the framework outperforms existing benchmark methods in terms of Signal-to-Noise Ratio (SNR), highlighting its advantage in semantic communication quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。