研究如何从粗粒度音频编码令牌更好重建波形,提升语音生成质量。
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
- 对比令牌预测与回归两种学习目标,优化重建策略
- 提出基于薛定谔桥的新方法,显著改善音频保真度
- 兼顾机器与人耳感知,适合语音合成与音频重建场景
神经音频编码器最初作为压缩技术设计,近年来在语音生成中受到关注。编码模型将每个音频帧表示为一系列离散令牌(即离散嵌入),这些令牌以不同粒度层次编码信息,从粗略到精细。现有工作多聚焦于如何更好生成粗粒度令牌,而本文关注一个同样重要但常被忽视的问题:如何更优地从粗粒度令牌重建波形?我们指出,学习目标和重建方法的选择对生成音频质量有显著影响。具体而言,研究了基于令牌预测与回归的两种策略,并引入一种基于薛定谔桥(Schrödinger Bridge)的新方法。通过分析不同设计选择对机器与人类感知的影响,验证了所提方法在音质上的优势。
原文摘要 · Abstract (English)
Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. The discrete and low-frequency nature of neural codecs introduced a new way to generate speech with token-based models. As these tokens encode information at various levels of granularity, from coarse to fine, most existing works focus on how to better generate the coarse tokens. In this paper, we focus on an equally important but often overlooked question: How can we better resynthesize the waveform from coarse tokens? We point out that both the choice of learning target and resynthesis approach have a dramatic impact on the generated audio quality. Specifically, we study two different strategies based on token prediction and regression, and introduce a new method based on Schrödinger Bridge. We examine how different design choices affect machine and human perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。