开源平台ESPnet-Codec统一训练评估音频神经编码器,支持语音音乐多场景应用。
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
- 基于ESPnet构建,提供语音、音乐、音频的编码器训练与评测配方。
- 集成6个任务,支持20项指标的全面性能评估。
- 适合研究音频生成、语音处理及模型对比的开发者使用。
神经编码器已成为近期语音和音频生成研究的关键。除了信号压缩能力外,离散编码器还被发现能提升下游训练效率,并与自回归语言模型兼容。然而,随着下游应用不断拓展,确保不同应用场景间的公平比较面临挑战。为此,我们提出全新的开源平台ESPnet-Codec,基于ESPnet构建,专注于神经编码器的训练与评估。ESPnet-Codec 提供涵盖语音、音乐和音频的多种训练与评估配方,支持多个广泛使用的编码器模型。同时,我们推出 VERSA——一个独立的评估工具包,可对编码器性能进行20项音频评估指标的全面分析。值得注意的是,ESPnet-Codec 可无缝集成至六个ESPnet任务中,支持多样化的实际应用。
原文摘要 · Abstract (English)
Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with autoregressive language models. However, as extensive downstream applications are investigated, challenges have arisen in ensuring fair comparisons across diverse applications. To address these issues, we present a new open-source platform ESPnet-Codec, which is built on ESPnet and focuses on neural codec training and evaluation. ESPnet-Codec offers various recipes in audio, music, and speech for training and evaluation using several widely adopted codec models. Together with ESPnet-Codec, we present VERSA, a standalone evaluation toolkit, which provides a comprehensive evaluation of codec performance over 20 audio evaluation metrics. Notably, we demonstrate that ESPnet-Codec can be integrated into six ESPnet tasks, supporting diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。