arXiv:2606.10591cs.SD2026-06中稿 · Interspeech 2026被引 1

让语音编码在极低码率下仍保持清晰可懂,核心是用内容特征引导重建。

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding

论文配图:ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding
图 1 · 摘自论文原文
  • 分双分支编码:一个专注声学细节,一个专传语言内容特征。
  • 在500 bps下仍保持高可懂度和听感质量,推理速度达0.4886 RTF。
  • 适合资源受限场景,如移动端超低码率语音通信。

神经语音编码器可在低码率下实现语音通信,但在超低码率(<1000 bps)下保持感知质量和可懂度仍具挑战。现有方法多侧重声学细节,导致在严格码率约束下核心语言信息容量不足。为此,我们提出ContextCodec,通过传输聚焦内容的上下文特征来显式引导重建。该模型采用双分支编码结构,将声学细节与内容导向的上下文特征解耦。上下文分支使用类似CLIP的对比损失训练,使上下文特征与音素索引对齐,减少副语言信息泄露。解码时,这些特征在每一步解码阶段注入以提供明确指导。此外,引入轻量级自回归潜在细化模块。实验表明,该方法在低至500 bps时仍保持优异的质量-可懂度权衡,在典型移动CPU上实现实时因子(RTF)为0.4886。

原文摘要 · Abstract (English)

Neural speech codecs enable low-bitrate speech communication, yet at ultra-low bitrates (< 1000 bps) preserving perceptual quality and intelligibility is challenging. Existing designs often prioritize acoustic details, leaving limited capacity for the core linguistic message under tight bitrate constraints. To address this, we propose ContextCodec, a codec that transmits content-focused context features to explicitly guide reconstruction. ContextCodec adopts a dual-branch encoder that decouples acoustic details from content-focused context. The context branch is trained with a CLIP-style contrastive loss that aligns context features with phoneme indices, reducing paralinguistic leakage. During decoding, these features are injected at each decoding stage for explicit guidance. In addition, we introduce a lightweight autoregressive latent refinement module. Experiments show a strong quality-intelligibility trade-off down to 500 bps, with an RTF of 0.4886 on a typical mobile CPU.

语音编码低码率内容引导神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。