评测语音助手在真实场景下的多轮交互能力,发现语音表现仅为文本的30%-45%。
$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
- 构建真实语音用户模拟器,支持多样口音与环境噪声,实现全双工交互评测
- 278个任务测试显示,语音助手完成率仅31%-51%,低于文本模式30%-45%
- 适用于评估语音助手在复杂现实场景中的可靠性与自然性
全双工语音助手——能同时听和说的系统——正从研究走向实际应用。然而现有评估方法仅关注对话动态或任务完成率,彼此孤立。我们提出τ-voice,一个面向真实世界复杂任务的语音助手基准测试框架:要求代理在多轮对话中遵循领域规则、与环境交互并完成可验证任务。该框架将τ²-bench扩展为支持语音交互的新型评测体系,结合可验证的任务完成度、全双工交互与真实音频数据,实现语音与文本性能的直接对比。一个可控且逼真的语音用户模拟器提供多样化口音、真实音频环境及丰富的对话轮次动态;通过解耦模拟与实时时间,可使用最强大的大模型而无实时限制。我们在278项任务上评估任务完成率(pass@1)与语音交互质量:尽管GPT-5(推理版)达到85%完成率,语音助手在干净条件下仅达31%-51%,在含噪声和多样口音的真实条件下降至26%-38%,仅保留文本能力的30%-45%。定性分析表明,79%-90%的失败源于代理行为,说明在当前设置下观察到的失败主要反映代理自身表现。τ-voice为衡量语音助手向自然、流畅、可靠方向演进提供了可复现的测试平台。
原文摘要 · Abstract (English)
Full-duplex voice agents--systems that listen and speak simultaneously--are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We introduce $τ$-voice, a benchmark for evaluating voice agents on grounded tasks with real-world complexity: agents must navigate complex multi-turn conversations, adhere to domain policies, and interact with the environment. The framework extends $τ^2$-bench into a novel voice agent benchmark combining verifiable completion of complex grounded tasks, full-duplex interaction, and realistic audio--enabling direct comparison between voice and text performance. A controllable and realistic voice user simulator provides diverse accents, realistic audio environments, and rich turn-taking dynamics; by decoupling simulation from wall-clock time, the user simulator can use the most capable LLM without real-time constraints. We evaluate task completion (pass@1) and voice interaction quality across 278 tasks: while GPT-5 (reasoning) achieves 85%, voice agents reach only 31--51% under clean conditions and 26--38% under realistic conditions with noise and diverse accents--retaining only 30--45% of text capability; qualitative analysis confirms 79--90% of failures stem from agent behavior, suggesting that observed failures primarily reflect agent behavior under our evaluation setup. $τ$-voice provides a reproducible testbed for measuring progress toward voice agents that are natural, conversational, and reliable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。