新基准测试显示语音助手在日常任务中仍有巨大提升空间。
DuplexWorld: Can voice agents help you get through the day?

- 构建六类生活场景的语音交互测试集,覆盖156个真实任务
- 最佳模型在对话流畅性、任务完成率等指标上均未超0.7
- 揭示语音助手在路径规划等复杂任务中的失效模式
语音代理正被广泛应用于企业客服和消费者日常陪伴,因其对话模态比文本更自然。然而现有评估基准仅关注工具调用与数据库查询,未能全面衡量语音代理在真实日常活动中的表现。为此,我们提出DuplexWorld,涵盖银行、保险、旅行、医疗、物流和路径规划六类应用场景,包含156个具体场景(350+小时对话),测试十一类对话任务。通过综合评估代理能力、对话连贯性及语音自然度,发现即使最优模型在任务成功率(Pass@1: 0.490)、对话轮次衔接(turn-taking: 0.653)和语音质量(DNSMOS: 3.378)方面仍存在明显不足。进一步分析揭示了代理在不同世界及对话类型中的表现差异,以及路径规划任务中探索与利用策略的失衡问题,并评估了各场景下的可靠性表现。
原文摘要 · Abstract (English)
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。