评测全双工对话系统在多轮交互中的表现,发现其易混淆、难纠错、易丢上下文。
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
- 设计带自动评估员的流式评测框架,分快慢节奏测试多轮对话能力。
- 实测表明全双工系统在并发说话时常混乱,纠错和实体追踪表现差。
- 开源协议与任务集,便于社区扩展新场景,加速全双工系统评估。
尽管全双工语音代理通过同时听与说实现了自然、低延迟的交互,但其在多轮场景下的一致性与任务表现仍缺乏深入研究。我们提出全双工评测基准v2(FDB-v2),一个集成自动评估员的流式框架,在快慢两种节奏设置下施加分阶段目标。FDB-v2涵盖四大任务类型:日常事务、修正指令、实体追踪与安全任务。评估指标包括对话轮次流畅性、多轮指令遵循能力及任务特定胜任力。该框架可扩展支持商业API与开源模型。测试结果显示,全双工系统在双方同时说话时容易混淆,难以平滑处理修正,且时常丢失对话主体或对象。通过开源的标准化流式协议与任务集,FDB-v2使新任务类别的扩展变得简便,助力社区加速对多轮全双工系统的评估。
原文摘要 · Abstract (English)
While full-duplex speech agents enable natural, low-latency interaction by speaking and listening simultaneously, their consistency and task performance in multi-turn settings remain underexplored. We introduce Full-Duplex-Bench-v2 (FDB-v2), a streaming framework that integrates with an automated examiner that enforces staged goals under two pacing setups (Fast vs. Slow). FDB-v2 covers four task families: daily, correction, entity tracking, and safety. We report turn-taking fluency, multi-turn instruction following, and task-specific competence. The framework is extensible, supporting both commercial APIs and open source models. When we test full-duplex systems with FDB-v2, they often get confused when people talk at the same time, struggle to handle corrections smoothly, and sometimes lose track of who or what is being talked about. Through an open-sourced, standardized streaming protocol and a task set, FDB-v2 makes it easy to extend to new task families, allowing the community to tailor and accelerate evaluation of multi-turn full-duplex systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。