arXiv:2506.04779cs.CLcs.SD2025-06被引 122

构建首个覆盖47类任务的语音理解与推理基准,推动人机语音交互发展。

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

论文配图:MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
图 1 · 摘自论文原文
  • 基于语音中的音韵、语调、情感等多维特征设计跨任务评测体系
  • 涵盖5000个音频-问题-答案对,覆盖语音理解全链条能力
  • 为语音大模型提供可量化的评估标准,助力研发更智能对话系统

语音蕴含远超文字的丰富声学信息。真实场景中的语音理解需融合语义内容、副语言特征(如情绪、语速、音高)和音系特征(如语调、重音、节奏)。尽管近期多模态语音大模型在处理音频信息方面展现出强大能力,其在自然语音中进行细粒度感知与复杂推理的能力仍待深入探索。为此,我们提出MMSU,一个专为语音理解与推理设计的综合性基准,包含5000个精心构建的音频-问题-答案三元组,覆盖47项不同任务。基于语言学理论,系统性地整合了语音学、语调、修辞、句法、语义及副语言等多种语言现象。通过对14个先进SpeechLLMs的严格评估,发现现有模型仍有显著提升空间,指明未来优化方向。MMSU为语音理解能力的全面评估树立新标准,为构建更高级的人机语音交互系统提供重要参考。基准数据集可在https://huggingface.co/datasets/ddwang2000/MMSU获取,评估代码见https://github.com/dingdongwang/MMSU。

原文摘要 · Abstract (English)

Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requires integrating semantic meaning (e.g., content), paralinguistic features (e.g., emotions, speed, pitch) and phonological characteristics (e.g., prosody, intonation, rhythm), which are embedded in speech. While recent multimodal Speech Large Language Models (SpeechLLMs) have demonstrated remarkable capabilities in processing audio information, their ability to perform fine-grained perception and complex reasoning in natural speech remains largely unexplored. To address this gap, we introduce MMSU, a comprehensive benchmark designed specifically for understanding and reasoning in spoken language. MMSU comprises 5,000 meticulously curated audio-question-answer triplets across 47 distinct tasks. To ground our benchmark in linguistic theory, we systematically incorporate a wide range of linguistic phenomena, including phonetics, prosody, rhetoric, syntactics, semantics, and paralinguistics. Through a rigorous evaluation of 14 advanced SpeechLLMs, we identify substantial room for improvement in existing models, highlighting meaningful directions for future optimization. MMSU establishes a new standard for comprehensive assessment of spoken language understanding, providing valuable insights for developing more sophisticated human-AI speech interaction systems. MMSU benchmark is available at https://huggingface.co/datasets/ddwang2000/MMSU. Evaluation Code is available at https://github.com/dingdongwang/MMSU.

语音理解多模态评测基准SpeechLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。