arXiv:2411.05361cs.CLeess.AS2024-11被引 96

180项任务的语音模型评测基准,助力通用语音模型能力评估

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

  • 构建180项任务的动态开放评测基准,涵盖语音、音乐与环境音频
  • 引入回归与序列生成等新任务,突破传统分类限制
  • 适合语音模型研发者、评测人员及跨模态研究者参考

多模态基础模型如Gemini和ChatGPT已显著提升人机交互体验。构建能理解多样自然语言指令的通用语音模型,对弥合沟通鸿沟至关重要。然而缺乏全面评测基准仍是关键挑战。我们提出Dynamic-SUPERB Phase-2,一个开源且持续扩展的通用语音模型评测基准。相较于第一代仅限分类任务,第二代新增125项由全球研究社区协作贡献的任务,总计达180项,成为目前最大的语音与音频评测基准。任务类型涵盖回归与序列生成,覆盖语音、音乐及环境音频。评测结果显示,无模型在所有任务上表现优异;SALMONN-13B在英文语音识别中表现最佳,Qwen2-Audio-7B-Instruct在情感识别中准确率最高。当前模型仍需创新以应对更广泛任务。所有任务数据与评估流程已开源至https://github.com/dynamic-superb/dynamic-superb。

原文摘要 · Abstract (English)

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication gaps and facilitating more intuitive interactions. However, the absence of a comprehensive evaluation benchmark poses a significant challenge. We present Dynamic-SUPERB Phase-2, an open and evolving benchmark for the comprehensive evaluation of instruction-based universal speech models. Building upon the first generation, this second version incorporates 125 new tasks contributed collaboratively by the global research community, expanding the benchmark to a total of 180 tasks, making it the largest benchmark for speech and audio evaluation. While the first generation of Dynamic-SUPERB was limited to classification tasks, Dynamic-SUPERB Phase-2 broadens its evaluation capabilities by introducing a wide array of novel and diverse tasks, including regression and sequence generation, across speech, music, and environmental audio. Evaluation results show that no model performed well universally. SALMONN-13B excelled in English ASR and Qwen2-Audio-7B-Instruct showed high accuracy in emotion recognition, but current models still require further innovations to handle a broader range of tasks. We open-source all task data and the evaluation pipeline at https://github.com/dynamic-superb/dynamic-superb.

语音模型评测基准多模态开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。