80 bps实现自然语音通信,分离文本、语调和音色分别压缩。
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
- 将语音分解为文本、语调、音色三部分,分别用不同方式高效压缩。
- 相比Opus降低75倍比特率,NISQA评分超4.26,抗丢包和噪声能力强。
- 模块化设计支持隐私加密与边缘部署,适合军事/卫星等低带宽场景。
在带宽受限的场景(如海事、卫星、战术网络)中,语音通信成本高昂。传统编码器在1 kbps以下表现不佳,现有语义方法(STT-TTS)则牺牲了语调和说话人身份。我们提出STCTS,一种生成式语义压缩框架,可在80 bps下实现自然语音通信。STCTS显式分解语音为语言内容、语调表达和说话人音色,分别采用针对性压缩策略:上下文感知文本编码(70 bps)、通过TTS插值稀疏传输语调(0.1–1 Hz时低于14 bps),以及分摊的说话人嵌入。在LibriSpeech上的评估显示,相比Opus(6 kbps)降低75倍比特率,相比EnCodec(1 kbps)降低12倍,同时保持感知质量(NISQA MOS > 4.26),具备良好的丢包退化特性和抗噪能力。我们还发现语调采样率存在双峰质量分布:稀疏与密集更新均达高质,中等速率因感知不连续而下降,为配置优化提供指导。除高效性外,其模块化架构支持隐私加密、可读传输与边缘设备灵活部署,为超低带宽场景提供稳健解决方案。
原文摘要 · Abstract (English)
Voice communication in bandwidth-constrained environments--maritime, satellite, and tactical networks--remains prohibitively expensive. Traditional codecs struggle below 1 kbps, while existing semantic approaches (STT-TTS) sacrifice prosody and speaker identity. We present STCTS, a generative semantic compression framework enabling natural voice communication at 80 bps. STCTS explicitly decomposes speech into linguistic content, prosodic expression, and speaker timbre, applying tailored compression: context-aware text encoding (70 bps), sparse prosody transmission via TTS interpolation (<14 bps at 0.1-1 Hz), and amortized speaker embedding. Evaluations on LibriSpeech demonstrate a 75x bitrate reduction versus Opus (6 kbps) and 12x versus EnCodec (1 kbps), while maintaining perceptual quality (NISQA MOS > 4.26), graceful degradation under packet loss and noise resilience. We also discover a bimodal quality distribution with prosody sampling rate: sparse and dense updates both achieve high quality, while mid-range rates degrade due to perceptual discontinuities--guiding optimal configuration design. Beyond efficiency, our modular architecture supports privacy-preserving encryption, human-interpretable transmission, and flexible deployment on edge devices, offering a robust solution for ultra-low bandwidth scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。