评测语音模型在真实口语场景下的多步骤工具使用能力,发现纠错和复杂推理仍是主要瓶颈。
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
- 构建真实人类语音数据集,含五类口误标注,支持跨任务链式API调用
- GPT-Realtime在准确率和打断避免上表现最佳,但延迟较高;Gemini Live 3.1最快但对话率最低
- 自纠正处理与复杂场景推理是所有系统共性短板,适合评估实时语音助手
我们提出Full-Duplex-Bench-v3(FDB-v3),一个用于评估语音模型在自然语音条件下进行多步工具调用的基准。该数据集完全由真实人类音频构成,标注了五类口误,并匹配四类任务域中的链式API调用场景。我们评估了六种模型配置:GPT-Realtime、Gemini Live 2.5、Gemini Live 3.1、Grok、Ultravox v0.7,以及传统级联流水线(Whisper→GPT-4o→TTS),涵盖准确性、延迟和对话轮次维度。GPT-Realtime在Pass@1上表现最优(0.600),打断避免率达13.5%;Gemini Live 3.1延迟最短(4.25~s),但对话轮次率最低(78.0%);级联基线虽有100%轮次率,但延迟最高(10.12~s)。所有系统在自纠正处理及高难度场景下的多步推理中均表现出持续失败。
原文摘要 · Abstract (English)
We introduce Full-Duplex-Bench-v3 (FDB-v3), a benchmark for evaluating spoken language models under naturalistic speech conditions and multi-step tool use. Unlike prior work, our dataset consists entirely of real human audio annotated for five disfluency categories, paired with scenarios requiring chained API calls across four task domains. We evaluate six model configurations -- GPT-Realtime, Gemini Live 2.5, Gemini Live 3.1, Grok, Ultravox v0.7, and a traditional Cascaded pipeline (Whisper$\rightarrow$GPT-4o$\rightarrow$TTS) -- across accuracy, latency, and turn-taking dimensions. GPT-Realtime leads on Pass@1 (0.600) and interruption avoidance (13.5\%); Gemini Live 3.1 achieves the fastest latency (4.25~s) but the lowest turn-take rate (78.0\%); and the Cascaded baseline, despite a perfect turn-take rate, incurs the highest latency (10.12~s). Across all systems, self-correction handling and multi-step reasoning under hard scenarios remain the most consistent failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。