arXiv:2607.20460cs.CLcs.AI2026-07被引 1

测试语音对话系统能否听懂指令调整说话时机,发现效果普遍不理想。

Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?

论文配图:Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
图 1 · 摘自论文原文
  • 构建指令可控的合成对话数据集,评估系统是否能按指令调整发言节奏。
  • 顶尖模型仅64.4%遵循指令,主动插话、回应等行为仍难控制。
  • 为需要灵活对话策略的应用(如教学、咨询)提供关键评测基准。

当前全双工语音对话系统虽能实现流畅交互,但其是否能在明确指令下调整发言时机尚不清晰。这在真实场景中至关重要,因不同应用对对话策略要求不同(如主动教学与被动倾听)。本文提出Instruct-FD,一个基于指令的全双工对话可控发言管理评测基准。为此,我们开发了经人工验证的可扩展合成管道,生成带指令的对话数据,并设计了不依赖部署环境的多轮评估协议及基于大模型的评判机制。对六种主流全双工系统的测试显示,指令遵循能力存在显著差距,最佳模型仅达64.4%的指令遵从率。不同行为和场景间表现差异大,主动行为(如回话、打断)尤其难以控制。研究结果确立了指令引导的发言管理是构建可适应、可部署全双工系统的关键方向。

原文摘要 · Abstract (English)

Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed. This is critical for real-world deployment, where conversational policies vary across applications (e.g., proactive tutoring vs. passive counseling). We introduce Instruct-FD, an instruction-conditioned benchmark for evaluating controllable turn management in FD systems. To enable this, we develop a human-validated, scalable synthetic pipeline that generates instruction-conditioned conversations, along with a deployment-agnostic multi-turn evaluation protocol and an LLM-based judge. Benchmarking six state-of-the-art full-duplex systems reveals a substantial gap in instruction-following turn management: the best model achieves only 64.4% adherence. Performance is highly uneven across behaviors and scenarios, with proactive behaviors such as model backchanneling and interruption remaining particularly challenging. These findings establish instruction-following turn management as a crucial direction for building adaptable and deployable full-duplex dialogue systems.

对话系统全双工指令跟随语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。