arXiv:2604.26296eess.AS2026-04中稿 · ICME 2026被引 1

探索语义先验在极低码率语音编码中的边界,提升清晰度与抗噪能力。

SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding

  • 引入冻结的语义先验(HuBERT/Whisper)增强低码率下的语义一致性。
  • 发现语义约束在1.5kbps时可降10%词错误率,6kbps后效果快速衰减。
  • 提出按码率动态调节先验强度,兼顾语义准确与语音自然度。

传统神经语音编码器在极低码率下会出现严重可懂度下降,瓶颈从声学失真转为语义丢失。本文系统研究了冻结语义先验(HuBERT 和 Whisper)在神经语音编码中的作用与基本极限。我们首次揭示并量化了一种新型‘语义退化’现象:在1.5 kbps时,语义约束可将词错误率(WER)相对降低约10%,但超过6 kbps后其增益迅速衰减,表明存在实际容量边界。进一步发现不同先验类型存在明显权衡:声学丰富的先验(HuBERT)更利于保留语调与音色细节,而高层语言先验(Whisper)在噪声环境中显著抑制语音幻觉(幻觉率降低26%),并大幅缩小对未见说话人的泛化差距。基于此,我们提出一种码率感知的调节策略,动态调整先验强度以优化语义一致性和听觉自然度之间的平衡。大量实验验证,该方法在可懂度和噪声鲁棒性上均达到现有基线水平,为生成式极低码率语音编码提供了有原则的路径。

原文摘要 · Abstract (English)

Conventional neural speech codecs suffer from severe intelligibility degradation at ultra-low bitrates, where the bottleneck transitions from acoustic distortion to semantic loss. To address this issue, this paper conducts a systematic investigation into the role and fundamental limits of integrating frozen semantic priors -- specifically HuBERT and Whisper -- into neural speech coding. We introduce and quantitatively validate a novel Semantic Retirement phenomenon: while semantic constraints reduce the Word Error Rate (WER) by up to ~10% relatively at 1.5 kbps, their benefits rapidly diminish beyond 6 kbps, indicating a practical capacity boundary. We further uncover a clear trade-off between different prior types: acoustic-rich priors (HuBERT) better preserve prosodic and timbral details, whereas high-level linguistic priors (Whisper) effectively suppress phonetic hallucinations in noisy environments (reducing hallucination rates by 26 percent) and substantially narrow the generalization gap for unseen speakers. Building on these findings, we propose a bitrate-aware regulation strategy that dynamically adjusts prior strength to optimize the trade-off between semantic consistency and perceptual naturalness. Extensive experimental evaluations confirm that our approach achieves competitive intelligibility and noise robustness compared to existing baselines, offering a principled pathway toward ultra-low-bitrate generative speech coding.

语音编码语义先验低码率生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。