构建多场景音频编解码评估基准,全面测试音质与语义保真度。
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
- 设计跨四类数据域的综合评测集,覆盖复杂应用场景
- 揭示现有编解码器在多说话人、噪声等场景下的性能瓶颈
- 适合音频编码、语音处理与大模型融合方向的研究者使用
随着多模态大语言模型的发展,音频编解码器在将音频转换为离散标记方面日益重要,使音频可被文本型大模型处理。当前编解码器需同时捕捉声学与语义信息。随着其应用于语音大模型的多样化场景(如多人对话、背景噪声、丰富的副语言信息),对编码能力的要求愈发复杂。然而,现有编解码器的评估仍受限于简单指标和单一场景,缺乏针对复杂应用的评测基准,难以全面评估其在声学与语义层面的表现。为此,我们提出 CodecBench,一个涵盖四个数据领域的综合性评测数据集,从声学与语义双角度评估编解码器性能。通过该基准,旨在识别当前局限性,指明未来研究方向,并推动音频编解码技术的发展。代码已开源:https://github.com/RayYuki/CodecBench。
原文摘要 · Abstract (English)
With the rise of multimodal large language models (LLMs), audio codec plays an increasingly vital role in encoding audio into discrete tokens, enabling integration of audio into text-based LLMs. Current audio codec captures two types of information: acoustic and semantic. As audio codec is applied to diverse scenarios in speech language model , it needs to model increasingly complex information and adapt to varied contexts, such as scenarios with multiple speakers, background noise, or richer paralinguistic information. However, existing codec's own evaluation has been limited by simplistic metrics and scenarios, and existing benchmarks for audio codec are not designed for complex application scenarios, which limits the assessment performance on complex datasets for acoustic and semantic capabilities. We introduce CodecBench, a comprehensive evaluation dataset to assess audio codec performance from both acoustic and semantic perspectives across four data domains. Through this benchmark, we aim to identify current limitations, highlight future research directions, and foster advances in the development of audio codec. The codes are available at https://github.com/RayYuki/CodecBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。