arXiv:2505.16845eess.AScs.AI2025-05中稿 · Interspeech 2025被引 12

首次让语音编码器支持可变帧率,动态调整编码效率。

Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate

  • 引入可变帧率机制,根据语音时序熵动态调整编码帧率。
  • 在低帧率下仍保持良好重建质量,平均帧率可灵活调节。
  • 适合实时语音传输与低延迟下游任务,提升编码效率。

大多数神经语音编码器通过帧内机制(如码本丢弃)实现比特率调整,且固定帧率(CFR)。然而语音段具有时变信息密度(如静音段与有声区差异),导致固定帧率并非最优,影响实时应用中的比特率与序列长度效率。本文首次将可变帧率(VFR)引入神经语音编码器,提出时序灵活性编码(TFC)技术。TFC可无缝调节平均帧率,并根据时序熵动态分配帧率。实验表明,采用TFC的编码器在保持高质量重建的同时具备高灵活性,即使在低帧率下性能仍具竞争力。该方法有望与其他低帧率神经语音编码研究结合,推动更高效下游任务的发展。

原文摘要 · Abstract (English)

Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently have time-varying information density (e.g., silent intervals versus voiced regions). This property makes CFR not optimal in terms of bitrate and token sequence length, hindering efficiency in real-time applications. In this work, we propose a Temporally Flexible Coding (TFC) technique, introducing variable frame rate (VFR) into neural speech codecs for the first time. TFC enables seamlessly tunable average frame rates and dynamically allocates frame rates based on temporal entropy. Experimental results show that a codec with TFC achieves optimal reconstruction quality with high flexibility, and maintains competitive performance even at lower frame rates. Our approach is promising for the integration with other efforts to develop low-frame-rate neural speech codecs for more efficient downstream tasks.

语音编码可变帧率神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。