arXiv:2508.08155eess.AScs.SD2025-08被引 11

构建多说话人对话理解基准,揭示模型在复杂场景下的能力短板

MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios

  • 分四层设计,从单人静态到多人交互逐步提升任务难度
  • 所有模型在高阶任务中性能显著下降,开源与闭源模型差距明显
  • 适合研究真实对话场景下语音理解的学者和工程师参考

语音理解(SLU)已从传统单任务方法发展为大型音频语言模型(LALM)解决方案。然而,现有语音基准大多聚焦单说话人或孤立任务,忽略了真实场景中常见的多说话人对话挑战。我们提出MSU-Bench,一个面向多说话人对话理解的综合性评估基准,采用以说话人为中心的设计。其层级框架涵盖四个渐进层次:单说话人静态属性理解、单说话人动态属性理解、多说话人背景理解以及多说话人交互理解。该结构确保所有任务均基于说话人中心语境,从基础感知延伸至跨说话人的复杂推理。在MSU-Bench上评估主流模型发现,随着任务复杂度上升,所有模型性能显著下降。同时观察到开源模型与闭源商业模型之间存在持续的能力差距,尤其在多说话人交互推理方面。这些结果验证了MSU-Bench在评估和推动真实多说话人环境对话理解方面的有效性。演示视频见附录。

原文摘要 · Abstract (English)

Spoken Language Understanding (SLU) has progressed from traditional single-task methods to large audio language model (LALM) solutions. Yet, most existing speech benchmarks focus on single-speaker or isolated tasks, overlooking the challenges posed by multi-speaker conversations that are common in real-world scenarios. We introduce MSU-Bench, a comprehensive benchmark for evaluating multi-speaker conversational understanding with a speaker-centric design. Our hierarchical framework covers four progressive tiers: single-speaker static attribute understanding, single-speaker dynamic attribute understanding, multi-speaker background understanding, and multi-speaker interaction understanding. This structure ensures all tasks are grounded in speaker-centric contexts, from basic perception to complex reasoning across multiple speakers. By evaluating state-of-the-art models on MSU-Bench, we demonstrate that as task complexity increases across the benchmark's tiers, all models exhibit a significant performance decline. We also observe a persistent capability gap between open-source models and closed-source commercial ones, particularly in multi-speaker interaction reasoning. These findings validate the effectiveness of MSU-Bench for assessing and advancing conversational understanding in realistic multi-speaker environments. Demos can be found in the supplementary material.

语音理解多说话人基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。