打造统一音频模型评估框架,支持多语言生成与编码器全面评测。
UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
- 构建模块化统一框架,集成24个主流模型与36个基准数据集。
- 提出三维度音频编码器评估法,覆盖语义、音色与声学质量。
- 新增中文语音评测集,助力中文语音模型客观评估。
自GPT-4o问世以来,音频基础模型发展迅速,但缺乏全面评估已成为制约进步的关键瓶颈,尤其在音频生成领域。当前音频评估面临三大挑战:(1)缺乏统一框架,数据集与代码分散,难以实现公平高效的跨模型对比;(2)音频编解码器缺少广泛认可的综合评估方法;(3)现有语音基准严重依赖英文,难以客观评估模型对中文的表现。为此,我们提出UltraEval-Audio,一个专为音频理解与生成任务设计的统一评估框架。该框架具备模块化架构,支持10种语言和14类核心任务,无缝整合24个主流模型与36个权威基准。为提升研究效率,框架提供一键式评估功能及实时公开排行榜。针对编码器评估,提出涵盖语义准确性、音色保真度和声学质量的三维评估方案。针对中文评估难题,构建SpeechCMMLU与SpeechHSK两个新基准,用于测评中文知识掌握与语言流利度。我们期望UltraEval-Audio能为学术界与产业界提供透明、高效、公平的音频模型对比平台。代码、基准与排行榜已开源:https://github.com/OpenBMB/UltraEval-Audio。
原文摘要 · Abstract (English)
The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particularly in audio generation. Current audio evaluation faces three major challenges: (1) audio evaluation lacks a unified framework, with datasets and code scattered across various sources, hindering fair and efficient cross-model comparison;(2) audio codecs, as a key component of audio foundation models, lack a widely accepted and holistic evaluation methodology; (3) existing speech benchmarks are heavily reliant on English, making it challenging to objectively assess models' performance on Chinese. To address the first issue, we introduce UltraEval-Audio, a unified evaluation framework for audio foundation models, specifically designed for both audio understanding and generation tasks. UltraEval-Audio features a modular architecture, supporting 10 languages and 14 core task categories, while seamlessly integrating 24 mainstream models and 36 authoritative benchmarks. To enhance research efficiency, the framework provides a one-command evaluation feature, accompanied by real-time public leaderboards. For the second challenge, UltraEval-Audio adopts a novel comprehensive evaluation scheme for audio codecs, evaluating performance across three key dimensions: semantic accuracy, timbre fidelity, and acoustic quality. To address the third issue, we propose two new Chinese benchmarks, SpeechCMMLU and SpeechHSK, designed to assess Chinese knowledge proficiency and language fluency. We wish that UltraEval-Audio will provide both academia and industry with a transparent, efficient, and fair platform for comparison of audio models. Our code, benchmarks, and leaderboards are available at https://github.com/OpenBMB/UltraEval-Audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。