arXiv:2512.20944cs.SD2025-12AAAI被引 1

1.5kbps下实现高保真语音编码,兼顾音质与语义。

SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs

  • 用非对称双量化器分离语义与声学特征,分别处理。
  • 1.5kbps下主观听感接近真实音频,下游任务语义更丰富。
  • 轻量级投影器注入语言先验,提升代码本利用率。

神经语音编码器在低比特率下面临保真度与语义丰富性之间的根本矛盾。为此,我们提出SACodec,一种基于非对称双量化器并引入语义锚定机制的新编码器。该设计通过轻量级投影器将声学特征对齐至冻结的大型mHuBERT代码本,注入语言先验并确保代码本充分使用,从而实现语义与声学特征的解耦量化。对于声学细节,采用带SimVQ的残差激活模块,使单层量化器能忠实恢复细微信息。在仅1.5 kbps的极低比特率下,SACodec在保真度和语义方面均达到新基准:主观听觉测试表明其重建质量与真实音频感知上高度相似,且在下游任务中展现出显著提升的语义丰富性。

原文摘要 · Abstract (English)

Neural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strategically decouples the quantization of Semantic and Acoustic details. The semantic anchoring is achieved via a lightweight projector that aligns acoustic features with a frozen, large-scale mHuBERT codebook, injecting linguistic priors while guaranteeing full codebook utilization. Sequentially, for acoustic details, a residual activation module with SimVQ enables a single-layer quantizer (acoustic path) to faithfully recover fine-grained information. At just 1.5 kbps, SACodec establishes a new state of the art by excelling in both fidelity and semantics: subjective listening tests confirm that its reconstruction quality is perceptually highly comparable to ground-truth audio, while its tokens demonstrate substantially improved semantic richness in downstream tasks.

语音编码低比特率语义保留量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。