评测语音助手在角色设定下隐式指令的遵循能力,揭示其实时对话控制的短板。
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

- 构建8类角色1038个测试用例,评估角色隐含行为的遵循度。
- 基于角色设定时系统行为准确率下降9.7%至4.5%,显式指令更可靠。
- 适用于研究语音交互中角色驱动行为与冲突指令处理的团队。
全双工语音助手需实时决策何时倾听、回应、打断、处理重叠语音、主动发言或让位。现有基准多依赖显式指令测试,而实际部署常通过角色或人格推断行为。本文提出DuplexSpeechBench-IFEval(DSB-IFEval),评估真实对话中隐式指令遵循能力。该基准包含1,038个测试用例,覆盖8种助手角色,评估五种条件:默认行为、显式指令、角色隐含行为、角色+规则联合、指令冲突。使用确定性指令遵从得分(IAS)衡量实时话语权管理,以大模型判断的角色一致性得分(PAS)评估内容一致性。在六个实时语音系统中发现架构依赖的权衡:如F-Actor和PersonaPlex在仅凭角色设定时遵从度分别下降9.7%和4.5%;而GPT-Realtime、MiniCPM-o、Fun-Audio-Chat虽内容一致,但话语权行为不随指令变化,主动行为受限。即使系统能遵循冲突指令,仍难以在安全冲突下主动纠正。结果表明,从角色推断行为、适时执行及解决指令冲突仍是全双工语音助手的核心挑战。
原文摘要 · Abstract (English)
Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. We introduce DuplexSpeechBench-IFEval (DSB-IFEval) for evaluating implicit instruction-following in real-time spoken interaction. (DSB-IFEval) comprises 1,038 test cases spanning eight diverse assistant roles and evaluates five conditioning protocols for instruction-following: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona--rule conditioning, and instruction conflict. We measure real-time floor management using a deterministic Instruction Adherence Score (IAS) and persona-consistent content using LLM-judged Persona Adherence Score (PAS). Across six real-time speech systems, we find architecture-dependent trade-offs. Full duplex models like F-Actor and PersonaPlex are more sensitive to whether conversational behavior is stated explicitly or must be inferred from a persona, with adherence dropping by 9.7% and 4.5%, respectively, under persona-only conditioning. In contrast, GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly adhere to persona-consistent content, but their floor behavior does not adapt across explicit and persona-only instructions and remains constrained on several proactive actions. We further find that even if systems reliably follow conflicting directives to their prescribed persona, they still struggle to override them under safety conflict. These results show that inferring the behavior implied by a role, executing it at the appropriate conversational moment, and resolving competing instructions remain distinct challenges for full-duplex voice agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。