构建多轮多语言对话基准,评估全双工语音系统表现
M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

- 设计支持英日语的多轮多领域对话评测集
- 发现模型在不同语言和场景下表现差异显著
- 适合研究全双工对话系统性能与交互机制
全双工语音对话系统(FDSDS)可在说话时持续聆听,实现自然的对话行为,如流畅换言、回应确认和用户打断处理。然而,多轮对话中的公平比较仍具挑战性,且现有基准对语言和对话领域的覆盖有限。本文提出 M3-DuplexBench,一个支持多轮、多语言、多领域的全双工对话评测基准。该基准涵盖英语与日语,覆盖日常对话与多轮问答任务。同时,评估模型在单轮、仅用户输入、教师强制全上下文等不同对话上下文设置下的表现,分析对话历史对模型行为的影响。对近期 FDSDS 模型的实验显示,模型存在特定的换言特征,跨语言与跨领域性能差距明显,且对话上下文影响效果复杂。
原文摘要 · Abstract (English)
Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。