评测语音对话系统在自然多轮交互中的真实表现,发现顶尖模型仍存严重缺陷。
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- 基于真实对话数据构建多轮语音评测框架,新增语音修复与环境线索测试
- 顶级模型在新基准上仅54.65%通过率,长上下文下自洽性显著下降
- 适合关注语音交互鲁棒性、自然对话能力的研究者与开发者
端到端语音对话系统正逐步取代传统流水线,直接处理原始音频。现有评测多基于合成语音和单轮任务,对真实多轮对话能力关注不足。本文提出 Audio MultiChallenge,一个开源基准,用于评估端到端语音对话系统在自然多轮交互中的表现。在文本型 MultiChallenge 框架基础上,新增语音编辑(Voice Editing)维度,测试模型对中途修正和回溯的鲁棒性;并扩展各维度至音频模态,如引入音频线索(Audio-Cue)挑战,要求模型记忆背景音与副语言信号。通过混合音频原生智能体与人工协作的流程,从47位说话人处收集452段对话,包含1,712条特定实例评判标准,保留真实语流中的不连贯性。对专有及开源模型的评估显示,即使最先进模型(Gemini 3 Pro Preview)也仅达54.65%通过率。错误分析表明,模型在新增维度上失败最多,且自洽性随音频上下文变长而恶化,反映出追踪语音修改、音频线索和长程依赖的困难。该基准为量化音频原生多轮交互能力提供可复现测试平台,推动技术改进。
原文摘要 · Abstract (English)
End-to-end (E2E) spoken dialogue systems are increasingly replacing cascaded pipelines for voice-based human-AI interaction, processing raw audio directly without intermediate transcription. Existing benchmarks primarily evaluate these models on synthetic speech and single-turn tasks, leaving realistic multi-turn conversational ability underexplored. We introduce Audio MultiChallenge, an open-source benchmark to evaluate E2E spoken dialogue systems under natural multi-turn interaction patterns. Building on the text-based MultiChallenge framework, which evaluates Inference Memory, Instruction Retention, and Self Coherence, we introduce a new axis Voice Editing that tests robustness to mid-utterance speech repairs and backtracking. We further augment each axis to the audio modality, such as introducing Audio-Cue challenges for Inference Memory that require recalling ambient sounds and paralinguistic signals beyond semantic content. We curate 452 conversations from 47 speakers with 1,712 instance-specific rubrics through a hybrid audio-native agentic and human-in-the-loop pipeline that exposes model failures at scale while preserving natural disfluencies found in unscripted human speech. Our evaluation of proprietary and open-source models reveals that even frontier models struggle on our benchmark, with Gemini 3 Pro Preview (Thinking), our highest-performing model achieving a 54.65% pass rate. Error analysis shows that models fail most often on our new axes and that Self Coherence degrades with longer audio context. These failures reflect difficulty of tracking edits, audio cues, and long-range context in natural spoken dialogue. Audio MultiChallenge provides a reproducible testbed to quantify them and drive improvements in audio-native multi-turn interaction capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。