arXiv:2501.05859eess.AS2025-01被引 3

用大模型实现多语言语音流式语义通信,低延迟高准确

Large Model Empowered Streaming Speech Semantic Communications

  • 将语义提取与编码移至边缘服务器,减轻本地设备负担
  • 支持多语言统一语义特征学习,传输准确率更高
  • 动态分段算法降低延迟,适合实时语音通信场景

本文提出一种基于大模型的流式语音语义通信系统LSSC-ST,支持跨语言语音传输。通过将复杂的语义提取与信道编码模块迁移至边缘服务器,降低本地设备计算压力。利用预训练大语音模型从多语言语音中学习统一语义特征,突破单一语言限制,提升实用性。输入语音以短片段形式逐段流式输入,配合新型动态语音分段算法自适应调整片段时长,进一步降低传输延迟。仿真结果表明,相比现有非流式语义通信系统,LSSC-ST在保持高质量输出的同时,实现了更低延迟的流式传输。

原文摘要 · Abstract (English)

In this paper, we introduce a large model-empowered streaming semantic communication system for speech transmission across various languages, named LSSC-ST. Specifically, we devise an edge-device collaborative semantic communication architecture by offloading the intricate semantic extraction and channel coding modules to edge servers, thereby reducing the computational burden on local devices. To support multilingual speech transmission, pre-trained large speech models are utilized to learn unified semantic features from speech in different languages, breaking the constraint of a single input language and enhancing the practicality of the LSSC-ST. Moreover, the input speech is sequentially streamed into the developed system as short speech segments, which enables low transmission latency without degrading the quality of the produced speech. A novel dynamic speech segmentation algorithm is proposed to further reduce the transmission latency by adaptively adjusting the duration of speech segments. According to simulation results, the LSSC-ST provides more accurate speech transmission and achieves a streaming manner with lower latency compared to the existing non-streaming semantic communication systems.

语音通信大模型流式传输多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。