用语义压缩技术实现超低带宽语音通信,质量不降反升。
A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication
- 将语音分解为高层语义特征,按需压缩不同信息。
- 在2-4倍更低码率下,语音识别等任务性能持平或超越传统编码器。
- 特别适合对语音语义敏感的下游任务,如说话人识别。
现有语音编码器虽能利用时间冗余并支持多尺度表示,但对音频所有特征一视同仁。而近期的生成式语音模型在文本转语音和语音迁移任务中已证明可将音频信号分解为语义上独立的高层表征。本文提出一种新型语义通信方法,利用此类表征,在极低码率下保持感知质量与下游任务适用性。实验表明,该方法在2-4倍更低码率下,于语音转录、情感分析和说话人验证任务中表现匹配或优于现有编码器;尤其在感知质量和说话人验证性能上显著超越Encodec,同时使用高达4倍更少的比特率。
原文摘要 · Abstract (English)
While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice models designed for text-to-speech and voice transfer tasks have recently proved effective at factorizing audio signals into high-level semantic representations of fundamentally distinct features. In this paper, we leverage such representations in a novel semantic communications approach to achieve lower bitrates without sacrificing perceptual quality or suitability for specific downstream tasks. Our technique matches or outperforms existing audio codecs on transcription, sentiment analysis, and speaker verification when encoding at 2-4x lower bitrate -- notably surpassing Encodec in perceptual quality and speaker verification while using up to 4x less bitrate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。