arXiv:2607.29125cs.CL2026-07

构建多轮多语言对话基准,评估全双工语音系统表现

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

论文配图:M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
图 1 · 摘自论文原文
  • 设计支持英日语的多轮多领域对话评测集
  • 发现模型在不同语言和场景下表现差异显著
  • 适合研究全双工对话系统性能与交互机制

全双工语音对话系统(FDSDS)可在说话时持续聆听,实现自然的对话行为,如流畅换言、回应确认和用户打断处理。然而,多轮对话中的公平比较仍具挑战性,且现有基准对语言和对话领域的覆盖有限。本文提出 M3-DuplexBench,一个支持多轮、多语言、多领域的全双工对话评测基准。该基准涵盖英语与日语,覆盖日常对话与多轮问答任务。同时,评估模型在单轮、仅用户输入、教师强制全上下文等不同对话上下文设置下的表现,分析对话历史对模型行为的影响。对近期 FDSDS 模型的实验显示,模型存在特定的换言特征,跨语言与跨领域性能差距明显,且对话上下文影响效果复杂。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.

对话系统多语言全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。