arXiv:2606.22868eess.AS2026-06中稿 · interspeech 2026被引 1

构建首个多说话人对话中的发言人中心理解评测基准。

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios

论文配图:MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
图 1 · 摘自论文原文
  • 设计两级框架,覆盖16类发言人相关任务,涵盖指代定位到对话推理。
  • 生成2300个问答对,人工验证确保答案有效性与标注一致性高。
  • 揭示模型在复杂发言人指代和多人对话推理中的系统性短板,适合评估语音大模型能力。

口语理解正从任务专用流程转向能生成自然语言回复的大型音频语言模型(LALMs)。然而,现有语音评测数据集主要关注单说话人场景或孤立子任务,未能充分评估真实多说话人对话中的发言人中心理解。本文提出 MSU-Bench,一个诊断性多说话人对话理解基准,包含16类发言人中心任务和2300个问答实例,采用两级框架,从发言人定位延伸至对话推理。我们构建了基于 Gemini 的辅助标注与问答生成流程,并通过人机协同验证,确保高问答有效性和人类答案与标注间强一致性。进一步分析发言人指代方案及诊断错误类型,揭示发言人定位与推理中的瓶颈。实验显示不同模型家族存在明显差距,闭源系统整体领先,但所有模型在复杂发言人定位和多说话人推理上仍面临挑战。基准数据、元信息及评估脚本将开源至 GitHub:https://github.com/ASLP-lab/MSU-Bench。

原文摘要 · Abstract (English)

Spoken Language Understanding (SLU) is moving from task-specific pipelines toward large audio language models (LALMs) that generate natural-language responses. However, existing speech benchmarks mainly focus on single-speaker settings or isolated subtasks, leaving speaker-centric understanding in realistic multi-speaker conversations insufficiently evaluated. We introduce MSU-Bench, a diagnostic benchmark for multi-speaker conversational understanding, covering 16 speaker-centric tasks and 2,300 QA instances in a two-tier framework from speaker grounding to dialogue reasoning. We build a Gemini-assisted annotation and QA generation pipeline with human-in-the-loop verification, achieving high QA validity and strong agreement between human answers and verified labels. We further analyze speaker-referencing schemes and diagnostic error types to reveal bottlenecks in speaker grounding and reasoning. Experiments reveal clear gaps across model families, with closed-source systems leading overall but all models still facing challenges in complex speaker grounding and multi-speaker reasoning. The benchmark annotations, metadata, and evaluation scripts will be available at the GitHub repository: https://github.com/ASLP-lab/MSU-Bench.

多说话人对话理解语音大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。