arXiv:2507.23159eess.AS2025-07中稿 · ICASSP 2026被引 42

首个自动化评估语音重叠处理的基准,揭示对话模型两种不同应对策略。

Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models

  • 构建四类重叠场景,自动测试模型在语音重叠时的表现。
  • 五款先进模型展现快速响应与保持对话流两种截然不同策略。
  • 开源工具助力开发者评测和优化全双工对话系统。

全双工语音对话系统有望将人机交互从僵化的轮流模式转变为自然流畅的对话。然而,管理语音重叠这一核心挑战仍严重缺乏评估。我们提出 Full-Duplex-Bench v1.5,首个完全自动化的基准,系统性探测模型在语音重叠下的行为。该基准模拟四种典型重叠场景:用户打断、用户回应词、与他人交谈及背景语音。框架兼容开源与商业API模型,提供涵盖对话行为分类、停止与响应延迟、语调适应等多维度指标。对五款最先进对话代理的评测揭示两种截然不同的策略:一种是优先快速响应用户输入的响应型策略,另一种是通过过滤重叠事件来维持对话流的主导型策略。我们的开源框架使从业者能够通过可复现的评估加速鲁棒全双工系统的开发。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue systems promise to transform human-machine interaction from a rigid, turn-based protocol into a fluid, natural conversation. However, the central challenge to realizing this vision, managing overlapping speech, remains critically under-evaluated. We introduce Full-Duplex-Bench v1.5, the first fully automated benchmark designed to systematically probe how models behave during speech overlap. The benchmark simulates four representative overlap scenarios: user interruption, user backchannel, talking to others, and background speech. Our framework, compatible with open-source and commercial API-based models, provides a comprehensive suite of metrics analyzing categorical dialogue behaviors, stop and response latency, and prosodic adaptation. Benchmarking five state-of-the-art agents reveals two divergent strategies: a responsive approach prioritizing rapid response to user input, and a floor-holding approach that preserves conversational flow by filtering overlapping events. Our open-source framework enables practitioners to accelerate the development of robust full-duplex systems by providing the tools for reproducible evaluation.

语音重叠对话系统评估基准全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。