动态遮蔽令牌,让语音编码更高效。
DTM-Codec: Dynamic Token Masking for VFR Speech Coding with Efficient Boundary Selection

- 根据语音变化动态保留关键帧,其余位置用学习的<MASK>填充
- 在相同总码率下,重建质量和可懂度显著优于固定帧率模型
- 新边界选择算法几乎无开销,适合实际语音压缩场景
变帧率(VFR)编码近年在神经语音编解码器中兴起,通过减少冗余区域的帧数、增加快速变化区域的帧数来提升效率。但VFR需传输保留时间步的附加信息,以往增益常因这些开销而被削弱。本文提出动态令牌遮蔽(DTM)-Codec,一种在严格匹配总码率条件下明显优于固定帧率基线的神经语音编解码器。DTM保留选定编码器令牌,将被遮蔽位置用学习的<MASK>嵌入填充,并传输二进制保留掩码以支持位置感知解码。我们进一步引入路径长度均衡(PLE)——一种线性时间边界选择器,可生成分布均匀的自适应段,开销极小。在多个码率点上,DTM-Codec普遍提升了重建质量与可懂度。
原文摘要 · Abstract (English)
Variable frame rate (VFR) coding has recently emerged in neural speech codecs, allocating fewer frames to redundant regions and more frames to rapidly changing speech. VFR must transmit side information about retained time steps, but prior gains are either not rigorously addressed or often minor once these overhead bits are included in total bitrate. We present Dynamic Token Masking (DTM)-Codec, a neural speech codec that demonstrates clear gains over fixed-frame-rate baselines under a strict matched-total-bitrate protocol. DTM keeps selected encoder tokens, fills masked positions with a learned <MASK> embedding, and transmits a binary keep-mask for position-aware decoding. We further introduce Path Length Equalization (PLE), a linear-time boundary selector for VFR coding that yields well-spread adaptive segments with negligible overhead. Across operating points, DTM-Codec broadly improves reconstruction quality and intelligibility over fixed-frame-rate baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。