arXiv:2512.04711cs.SDcs.AI2025-12被引 2

用大模型实现自适应语音语义通信,抗丢包能力强。

Large Speech Model Enabled Semantic Communication

  • 用大模型生成离散语音符号,支持动态压缩与传输
  • 在550-2060 bps带宽下,高丢包率时语音质量更优
  • 适合实时语音通信系统,如远程医疗、车联网

现有语音语义通信系统多基于联合信源信道编码(JSCC)架构,性能受限于特定任务和数据集的模型结构。近期研究表明,大规模预训练生成模型在少量微调下即可在多种下游任务中表现优异。为此,我们提出基于大语音模型的语义通信系统(LargeSC),利用大模型中的丰富语义知识,实现损毁信道下的自适应传输。同时实现高效压缩与鲁棒传输仍具挑战,需权衡压缩效率、语音质量和延迟。本文采用Mimi作为语音编解码器,将语音转换为与现有网络兼容的离散符号。设计自适应控制器模块,支持带内不等错误保护(UEP),动态调节传输策略以应对语音内容和丢包概率变化。此外,使用低秩微调(LoRA)对Moshi基础模型进行微调,用于丢失语音符号的生成式恢复。仿真结果表明,该系统支持550至2060 bps的带宽范围,在高丢包率下语音质量优于传统基线,端到端延迟约460毫秒,具备实时部署潜力。

原文摘要 · Abstract (English)

Existing speech semantic communication systems mainly based on Joint Source-Channel Coding (JSCC) architectures have demonstrated impressive performance, but their effectiveness remains limited by model structures specifically designed for particular tasks and datasets. Recent advances indicate that generative large models pre-trained on massive datasets, can achieve outstanding performance arexhibit exceptional performance across diverse downstream tasks with minimal fine-tuning. To exploit the rich semantic knowledge embedded in large models and enable adaptive transmission over lossy channels, we propose a Large Speech Model enabled Semantic Communication (LargeSC) system. Simultaneously achieving adaptive compression and robust transmission over lossy channels remains challenging, requiring trade-offs among compression efficiency, speech quality, and latency. In this work, we employ the Mimi as a speech codec, converting speech into discrete tokens compatible with existing network architectures. We propose an adaptive controller module that enables adaptive transmission and in-band Unequal Error Protection (UEP), dynamically adjusting to both speech content and packet loss probability under bandwidth constraints. Additionally, we employ Low-Rank Adaptation (LoRA) to finetune the Moshi foundation model for generative recovery of lost speech tokens. Simulation results show that the proposed system supports bandwidths ranging from 550 bps to 2.06 kbps, outperforms conventional baselines in speech quality under high packet loss rates and achieves an end-to-end latency of approximately 460 ms, thereby demonstrating its potential for real-time deployment.

语音通信大模型语义通信抗丢包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。