构建轻量级评测基准,公平对比神经音频编码器性能。
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
- 设计统一评测框架,确保不同编码器公平比较。
- 选用免版权数据集并采样成小集,降低计算成本。
- 适合研究音频编码、语音建模的开发者参考。
神经音频编码器正日益重要,作为音频的分词器,支持高效传输或语音语言建模。理想的神经音频编码器应在低比特率下仍保持内容、语调、说话人特征和音频信息。近年来,大量先进编码模型被提出,但测试条件各异,难以横向对比。为此,我们在 SLT 2024 引入 Codec-SUPERB 挑战,旨在实现现有编码器的公平、轻量级比较,并推动领域发展。该挑战整合代表性语音应用与客观指标,精心选取免版权数据集,并将其采样为小型数据集以降低评估计算开销。本文介绍挑战规则、数据集、五种参与系统、评测结果及发现。
原文摘要 · Abstract (English)
Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 2024, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge's rules, datasets, five participant systems, results, and findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。