发现低帧率音频编码器质量下降主因是训练配置不当,非本质瓶颈。
Probing Low Frame Rate Degradation in Neural Audio Codecs
- 通过控制实验排除语音碰撞和码本饱和,定位问题根源
- 修正训练配置后,1.6赫兹仍可保持平滑降质,无突变断崖
- 对语音合成效率优化有实际价值,适合追求推理加速的研究者
神经音频编码器在低帧率下具有吸引力,因其生成成本与序列长度线性相关。近期研究已实现12.5赫兹及以下运行,但低帧率退化机制尚不明确。本文通过受控帧率消融实验,复现了此前报告的6.25赫兹质量断崖,并评估了语音单元冲突与码本饱和等候选解释,均未发现根本性障碍。断崖实为训练配置不当所致:训练时固定片段时长导致低帧率下令牌过少,使解码器缺乏跨令牌上下文。修正后,词错误率(WER)随音素负载平稳下降至3.1赫兹和1.6赫兹,表明低帧率编码器的推理效率优势比先前认为更易实现。
原文摘要 · Abstract (English)
Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost scales linearly with the sequence length. Recent work has demonstrated that codecs can operate at 12.5 Hz and below, but the mechanisms underlying low frame rate degradation remain insufficiently understood. We investigate these mechanisms through a controlled frame rate ablation. We reproduce a quality cliff at 6.25 Hz reported in previous works and evaluate candidate explanations: phonemic collisions and codebook saturation, neither of which shows evidence of a fundamental barrier. The cliff is instead caused by suboptimal training configuration: fixed clip duration during training yields too few tokens at low frame rates, starving the decoder of inter-token context. Once corrected, WER degrades smoothly with phonemic load down to 3.1 Hz and 1.6 Hz, suggesting the inference-time efficiency gains of low frame rate codecs are more accessible than previously assumed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。