让语音助手按角色精准抢话,减少误打断。
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
- 根据角色动态调整发言时机,用流式语音大模型实现
- 实测抢话准确率提升40%以上,漏判减少70%以上
- 适合会议助手、客服机器人等多人对话场景
多人群体语音对话中的发言权交接仍是语音代理的核心挑战,尤其在动态竞争和用户期望不同时。我们提出ModeratorLM,一种基于角色扮演的语音代理,其发言行为显式依赖于分配的角色。系统基于分块流式处理的语音大模型构建,并引入链式思维推理机制,结合对话上下文与角色信息进行决策。我们构建了大规模合成数据集RolePlayConv,包含多样化的助理角色。在真实会议数据和RolePlayConv上的实验表明,相比非角色条件基线,该方法将发言权交接准确率提升40%以上,召回率提升70%以上,且显著降低误打断率。
原文摘要 · Abstract (English)
Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations. We propose ModeratorLM, a role-playing voice agent that conditions turn-taking behavior on an explicitly assigned role in multi-party settings. The system is built on a speech large language model operating in chunk-wise streaming manner. We further introduce a reasoning-augmented variant that incorporates chain-of-thought reasoning over conversational context and the assigned role. We construct RolePlayConv, a large-scale synthetic dataset of spoken multi-party conversations with diverse assistant roles. Experiments on real-world meeting data and RolePlayConv show improved turn-taking precision by over 40% and recall by more than 70%, while substantially reducing false-positive interruptions compared to non-role-conditioned baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。