将X-Codec-2.0的隐状态率降至25Hz,采样率提至24kHz,提升语音压缩效率与音质。
Improving X-Codec-2.0 for Multi-Lingual Speech: 25 Hz Latent Rate and 24 kHz Sampling
- 通过增加池化层和扩大解码器步长,降低隐状态率至25Hz
- 在多语言Common Voice 17数据集上实现0.29分的感知质量提升(UTMOSv2)
- 无需改动核心结构,适合语音编码与跨语言语音处理研究者
X-Codec-2.0在神经音频压缩与多语言语音建模中表现优异,采用50 Hz隐状态率和16 kHz采样率,基于冻结的HuBERT特征。然而,该配置限制了时间效率与音频保真度。本文通过引入额外池化并增大解码器步长,将隐状态率从50 Hz降至25 Hz,同时将输出采样率从16 kHz提升至24 kHz,显著提高效率与感知质量,且未改变原有架构。在多语言Common Voice 17测试集上,新配置相比原版提升0.29 MOS(UTMOSv2),达到当前25 Hz速率下最佳性能。代码、模型权重及生成对比已公开于Hugging Face。
原文摘要 · Abstract (English)
X-Codec-2.0 has shown strong performance in neural audio compression and multilingual speech modeling, operating at a 50 Hz latent rate and a 16 kHz sampling rate using frozen HuBERT features. While effective, this configuration limits temporal efficiency and audio fidelity. In this work, we explore a simple and effective modification by introducing additional pooling and increasing the decoder hop size. This reduces the latent rate from 50 Hz to 25 Hz and simultaneously raises the output sampling rate from 16 kHz to 24 kHz, improving efficiency and perceptual quality without altering the core architecture. Evaluated on the multilingual Common Voice 17 test set, the proposed configuration achieves a 0.29 MOS improvement over the original X-Codec-2.0 baseline based on UTMOSv2, and attains the best reported performance among all codecs operating at 25 Hz. The source code, checkpoints, and generation comparisons are released at \href{https://huggingface.co/Scicom-intl/xcodec2-25TPS-24k}{https://huggingface.co/Scicom-intl/xcodec2-25TPS-24k}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。