用生成潜空间提升语音传输抗丢包能力,兼容现有系统。
Error-Resilient Semantic Communication for Speech Transmission over Packet-Loss Networks
- 在生成潜空间进行鲁棒语音编码,利用潜在先验实现高质量丢包隐藏。
- 相比传统前向纠错,减少冗余开销并提升动态网络下的鲁棒性。
- 适合需低延迟、高容错的实时语音通信场景,如远程会议与语音助手。
无线网络中实时语音通信仍面临挑战,传统信道保护机制在带宽和时延严格约束下难以有效应对丢包。语义通信通过联合源信道编码(JSCC)提升了语音传输的鲁棒性,但其跨层设计阻碍了与现有数字通信系统的兼容。为此,本文提出基于生成潜先验的鲁棒语音语义通信框架Glaris,可在生成潜空间实现抗丢包编码。生成潜先验使接收端能实现高质量丢包隐藏(PLC),在语义一致性与重建保真度间取得良好平衡。同时,集成的错误韧性机制有效缓解误差传播,提升PLC效果。实验在LibriSpeech数据集上显示,相较于传统包级前向纠错(FEC),Glaris显著增强动态无线网络下的鲁棒性,大幅降低冗余开销,达到接近JSCC水平的抗扰能力,且保持与现有系统的无缝兼容,兼顾传输效率与语音重建质量。
原文摘要 · Abstract (English)
Real-time speech communication over wireless networks remains challenging, as conventional channel protection mechanisms cannot effectively counter packet loss under stringent bandwidth and latency constraints. Semantic communication has emerged as a promising paradigm for enhancing the robustness of speech transmission by means of joint source-channel coding (JSCC). However, its cross-layer design hinders practical deployment due to the incompatibility with existing digital communication systems. In this case, the robustness of speech communication is consequently evaluated primarily by the error-resilience to packet loss over wireless networks. To address these challenges, we propose \emph{Glaris}, a generative latent-prior-based resilient speech semantic communication framework that performs resilient speech coding in the generative latent space. Generative latent priors enable high-quality packet loss concealment (PLC) at the receiver side, well-balancing semantic consistency and reconstruction fidelity. Additionally, an integrated error resilience mechanism is designed to mitigate the error propagation and improve the effectiveness of PLC. Compared with traditional packet-level forward error correction (FEC) strategies, our new method achieves enhanced robustness over dynamic wireless networks while reducing redundancy overhead significantly. Experimental results on the LibriSpeech dataset demonstrate that \emph{Glaris} consistently outperforms existing error-resilient codecs, achieving JSCC-level robustness while maintaining seamless compatibility with existing systems, and it also strikes a favorable balance between transmission efficiency and speech reconstruction quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。