arXiv:2508.18240cs.CLcs.AI2025-08被引 13

新基准MTalk-Bench揭示语音模型在多轮对话中的语义与非语言信息处理短板

MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols

  • 构建覆盖语义、副语言和环境音的三维度多轮对话评测框架
  • 模型在语义理解上表现好,但对语气和背景音感知差,长回复易失效率
  • 双评估法(对抗赛+评分表)更可靠,需标注辅助非言语评估

语音到语音大模型的快速发展显著提升了实时口语交互能力,但现有评估框架难以有效衡量复杂多轮对话中的表现。为此,我们提出MTalk-Bench,一个涵盖语义信息、副语言信息和环境音三个核心维度的多轮语音对话基准。每个维度包含九个真实场景及对应任务,用于评估推理等特定能力。采用双方法评估框架:对抗式评价(成对比较)和评分表评价(绝对打分),结合人工与大模型评估。实验结果表明:(1) 模型在语义信息处理上表现优异,但在副语言信息与环境音感知方面表现不足;(2) 增加回复长度可恢复连贯性,但牺牲多轮对话效率;(3) 针对任务设计的模态感知模型优于单纯规模扩展。评估框架方面:(1) 对抗与评分表给出一致且互补的排名,但仅当性能差距明显时才可靠;(2) 大模型作为裁判在差距明显或标准明确时与人类一致,但在位置和长度上存在偏差,非语言评估需文本标注才可靠。研究揭示当前语音评测的局限性,亟需更鲁棒、语音感知更强的评估体系。

原文摘要 · Abstract (English)

The rapid advancement of speech-to-speech (S2S) large language models (LLMs) has significantly improved real-time spoken interaction. However, current evaluation frameworks remain inadequate for assessing performance in complex, multi-turn dialogues. To address this, we introduce MTalk-Bench, a multi-turn S2S benchmark covering three core dimensions: Semantic Information, Paralinguistic Information, and Ambient Sound. Each dimension includes nine realistic scenarios, along with targeted tasks to assess specific capabilities such as reasoning. Our dual-method evaluation framework combines Arena-style evaluation (pairwise comparison) and Rubrics-based evaluation (absolute scoring) for relative and absolute assessment. The benchmark includes both model and human outputs, evaluated by human evaluators and LLMs. Experimental results reveal two sets of findings. Overall performance of S2S LLMs: (1) models excel at semantic information processing yet underperform on paralinguistic information and ambient sounds perception; (2) models typically regain coherence by increasing response length, sacrificing efficiency in multi-turn dialogues; (3) modality-aware, task-specific designs outperform brute scaling. Evaluation framework and reliability: (1) Arena and Rubrics yield consistent, complementary rankings, but reliable distinctions emerge only when performance gaps are large; (2) LLM-as-a-judge aligns with humans when gaps are clear or criteria explicit, but exhibits position and length biases and is reliable on nonverbal evaluation only with text annotations. These results highlight current limitations in S2S evaluation and the need for more robust, speech-aware assessment frameworks.

语音生成多轮对话评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。