arXiv:2511.10262cs.CLcs.AI2025-11ACL被引 13

首个面向全双工语音模型多轮对话的综合评估基准

MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models

  • 将连续对话切分为离散回合,实现逐轮评估
  • 覆盖对话质量、指令遵循、安全等多个维度
  • 揭示现有模型在多轮中表现不稳的缺陷

全双工语音语言模型(FD-SLMs)支持实时重叠对话,相比传统半双工模型更具动态交互性。然而现有基准主要聚焦单轮交互,忽视多轮对话的复杂性。多轮评估面临说话人轮次边界模糊、推理时上下文不一致等挑战,且多数基准仅关注对话特性,忽略其他关键方面。为此,我们提出MTR-DuplexBench,一个针对FD-SLMs的多轮综合评估基准。该基准不仅将连续全双工对话分割为离散回合以实现逐轮评估,还涵盖对话特征、对话质量、指令遵循与安全性等多个维度。实验表明,当前FD-SLMs在多轮和多维度下难以保持稳定性能,凸显本基准的必要性与有效性。代码与数据已公开于:https://github.com/ZhangHe0918/MTR-DuplexBench

原文摘要 · Abstract (English)

Full-Duplex Speech Language Models (FD-SLMs) enable real-time, overlapping conversational interactions, offering a more dynamic user experience compared to traditional half-duplex models. However, existing benchmarks primarily focus on evaluating single-round interactions, neglecting the complexities of multi-round communication. Evaluating FD-SLMs in multi-round settings poses significant challenges, including blurred turn boundaries in communication and context inconsistency during model inference. Also, existing benchmarks often focus solely on evaluating conversational features, neglecting other critical aspects. To address these gaps, we introduce MTR-DuplexBench, a novel benchmark designed for a comprehensive multi-round evaluation of FD-SLMs. MTR-DuplexBench not only segments continuous full-duplex dialogues into discrete turns for turn-by-turn assessment but also incorporates various evaluation aspects, including conversational features, dialogue quality, instruction following, and safety. Experimental results reveal that current FD-SLMs face difficulties in maintaining consistent performance across multiple rounds and evaluation dimensions, highlighting the necessity and effectiveness of our benchmark. Code and data are available at: https://github.com/ZhangHe0918/MTR-DuplexBench

语音模型对话评估多轮交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。