arXiv:2501.12696eess.AScs.SD2025-01被引 8

用大模型提升音频传输抗丢包能力,压缩与纠错一体。

SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling

  • 分层设计:先压缩音频,再用大模型预测并修复丢包数据。
  • 在多种信道下,音质和感知质量显著优于现有系统。
  • 适合对稳定性要求高的语音通信、远程会议等场景。

本文提出SoundSpring,一种兼具鲁棒性与兼容性的新型音频传输系统。不同于以往直接将音频映射为模拟信号的深度联合源信道编码(JSCC)方法,SoundSpring采用分层架构,将音频压缩与数字编码传输分离,并充分挖掘大语言模型(foundation models)在上下文中的预测能力。通过引入因果顺序掩码学习策略,单一模型在潜在特征空间中同时实现高效压缩与丢包恢复。实验表明,该框架在信号保真度与感知质量上均显著优于当前主流音频传输系统,验证了掩码学习语言模型作为上下文预测器的有效性。研究结果不仅支持SoundSpring在基于学习的音频通信系统中部署,也为未来语义化音频传输器的设计提供了新思路。

原文摘要 · Abstract (English)

In this paper, we propose "SoundSpring", a cutting-edge error-resilient audio transceiver that marries the robustness benefits of joint source-channel coding (JSCC) while also being compatible with current digital communication systems. Unlike recent deep JSCC transceivers, which learn to directly map audio signals to analog channel-input symbols via neural networks, our SoundSpring adopts the layered architecture that delineates audio compression from digital coded transmission, but it sufficiently exploits the impressive in-context predictive capabilities of large language (foundation) models. Integrated with the casual-order mask learning strategy, our single model operates on the latent feature domain and serve dual-functionalities: as efficient audio compressors at the transmitter and as effective mechanisms for packet loss concealment at the receiver. By jointly optimizing towards both audio compression efficiency and transmission error resiliency, we show that mask-learned language models are indeed powerful contextual predictors, and our dual-functional compression and concealment framework offers fresh perspectives on the application of foundation language models in audio communication. Through extensive experimental evaluations, we establish that SoundSpring apparently outperforms contemporary audio transmission systems in terms of signal fidelity metrics and perceptual quality scores. These new findings not only advocate for the practical deployment of SoundSpring in learning-based audio communication systems but also inspire the development of future audio semantic transceivers.

音频传输大模型抗丢包压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。