arXiv:2607.05250cs.CL2026-07中稿 · SPECOM 2026

提出分层时不变表示,提升语音编码的低延迟生成能力。

Streaming Neural Speech Codecs through Time-Invariant Representations

论文配图:Streaming Neural Speech Codecs through Time-Invariant Representations
图 1 · 摘自论文原文
  • 用时不变模块分离语音内容与说话人/环境信息
  • 多层互补特征提升重建质量与说话人相似度
  • 支持660ms块流式处理,性能几乎无损失

神经语音编码器在基于编码器的语音生成系统中日益重要。TiCodec通过时间不变表示提取(TIRE)模块,将时变语音内容与时不变信息解耦,可能减少帧级建模的信息量。本文研究TIRE表示所捕捉信息的本质及其在低延迟语音处理中的适用性。通过一系列探针任务分析编码器层的影响,发现中间层捕获互补的说话人和环境相关特征,但包含极少语言内容。进一步研究了多种TIRE训练的段落选择策略,表明跨文件采样能增强不变表示的鲁棒性。基于此,提出双层TIRE(Dual-TIRE)架构,利用不同编码层的互补性,显著提升语音重建质量和说话人相似度。最后,在连续660ms处理块的流式推理设置下评估TiCodec,结果表明流式操作可实现且重建性能无明显下降,凸显因子化神经编码表示在未来的低延迟语音生成系统中的潜力。

原文摘要 · Abstract (English)

Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized representation that separates time-varying speech content from time-invariant information through a Time-Invariant Representation Extraction (TIRE) module, potentially reducing the amount of information that must be modeled at the frame-level. In this work, we investigate the nature of the information captured by TIRE representations and their suitability for low-latency speech processing. Using a series of probing tasks, we analyze the influence of the encoder layer and show that intermediate layers capture complementary speaker- and environment-related information while containing little linguistic content. We further study several segment selection strategies for TIRE training and demonstrate that cross-file sampling improves the robustness of invariant representations. Based on these findings, we propose Dual-TIRE, a multi-level architecture that exploits the complementarity of different encoder layers and improves speech reconstruction quality and speaker similarity. Finally, we evaluate TiCodec in a streaming inference setting using successive 660ms processing blocks. Results show that streaming operation can be achieved without significant degradation in reconstruction performance, highlighting the potential of factorized neural codec representations for future low-latency speech generation systems.

语音编码流式处理时不变表征低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。