arXiv:2507.19040eess.AScs.CL2025-07中稿 · Interspeech 2025被引 18

为全双工语音对话系统设计新评测流程,解决中断响应难题

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems

  • 用大模型+语音技术构建自动化评测流水线
  • 在1200次中断下测试三款系统均表现不佳
  • 适合研究全双工交互与鲁棒性评估的团队

全双工语音对话系统(FDSDS)通过支持实时用户打断和回应,使人机交互更自然,优于传统基于轮流对话的系统。然而现有评测基准缺乏对全双工场景的评估指标,如用户打断时的表现。本文提出一个综合评测流水线,结合大语言模型、语音合成与语音识别技术,针对中断处理、延迟应对及噪声环境下的鲁棒性,引入多项新指标。我们使用超过40小时生成语音,对三个开源全双工系统(Moshi、Freeze-omni、VITA-1.5)进行了293次模拟对话与1200次中断测试。结果表明,所有模型在频繁打断和嘈杂条件下仍难以有效响应用户中断。相关演示、数据与代码将公开。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue systems (FDSDS) enable more natural human-machine interactions by allowing real-time user interruptions and backchanneling, compared to traditional SDS that rely on turn-taking. However, existing benchmarks lack metrics for FD scenes, e.g., evaluating model performance during user interruptions. In this paper, we present a comprehensive FD benchmarking pipeline utilizing LLMs, TTS, and ASR to address this gap. It assesses FDSDS's ability to handle user interruptions, manage delays, and maintain robustness in challenging scenarios with diverse novel metrics. We applied our benchmark to three open-source FDSDS (Moshi, Freeze-omni, and VITA-1.5) using over 40 hours of generated speech, with 293 simulated conversations and 1,200 interruptions. The results show that all models continue to face challenges, such as failing to respond to user interruptions, under frequent disruptions and noisy conditions. Demonstrations, data, and code will be released.

语音对话全双工评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。