用自监督学习降低语义通信训练成本,提升效率。
Multi-Modal Self-Supervised Semantic Communication
- 预训练阶段用自监督提取通用语义特征
- 在NYU Depth V2上训练开销减少且性能不降
- 适合资源受限的边缘智能系统
语义通信是一种新兴范式,通过深度学习提取并传输语义信息。现有研究多关注降低通信开销,却忽视动态无线环境下训练阶段的巨大通信成本。为此,我们提出一种多模态自监督语义通信系统,利用多模态自监督学习增强任务无关特征提取能力。该方法在预训练阶段采用自监督学习提取任务无关语义特征,再通过有监督微调适配下游任务。这种双阶段策略有效捕捉模态不变与模态特定特征,同时显著降低训练通信开销。在NYU Depth V2数据集上的实验表明,该方法在保持或超越现有监督学习性能的同时,大幅减少训练相关通信开销。结果证明了多模态自监督学习在语义通信中的优势,为更高效、可扩展的边缘推理系统铺平道路。
原文摘要 · Abstract (English)
Semantic communication is emerging as a promising paradigm that focuses on the extraction and transmission of semantic meanings using deep learning techniques. While current research primarily addresses the reduction of semantic communication overhead, it often overlooks the training phase, which can incur significant communication costs in dynamic wireless environments. To address this challenge, we propose a multi-modal semantic communication system that leverages multi-modal self-supervised learning to enhance task-agnostic feature extraction. The proposed approach employs self-supervised learning during the pre-training phase to extract task-agnostic semantic features, followed by supervised fine-tuning for downstream tasks. This dual-phase strategy effectively captures both modality-invariant and modality-specific features while minimizing training-related communication overhead. Experimental results on the NYU Depth V2 dataset demonstrate that the proposed method significantly reduces training-related communication overhead while maintaining or exceeding the performance of existing supervised learning approaches. The findings underscore the advantages of multi-modal self-supervised learning in semantic communication, paving the way for more efficient and scalable edge inference systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。