提出双阶段语音切换机制,提升多人对话中发言权转移的准确率。
Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation

- 分两步处理:先快速定位可能的发言结束点,再轻量验证是否换人。
- 在VoxConverse数据集上,换人检测准确率优于基线模型。
- 用扩散模型增强音频数据,进一步提升换人识别效果,适合多说话人系统研究者。
可靠的发言权转移对语音对话系统至关重要。然而,现有方法大多针对双人对话,在包含重叠和快速换人的真实多人语音场景中表现不佳。本文基于VoxConverse数据集研究多人对话中的发言权转移问题,提出一种纯音频的两阶段流水线:将何时触发发言结束与是否真正换人分离处理。快速触发器扫描音频,提出候选发言结束时间;轻量验证器仅在这些时刻判断是保持(Hold)还是切换(Shift),并支持下一位说话人预测。在完整多人设置及可控的两人前两名投影设置下报告结果。同时研究了基于扩散模型的、保留标签的背景音频混合数据增强策略。实验表明,相比基线,换人检测性能提升,且扩散增强带来进一步改进。
原文摘要 · Abstract (English)
Reliable turn-taking is essential for spoken dialogue systems. However, most existing methods are designed for two-speaker interaction and struggle with realistic multiparty audio containing overlap and rapid speaker changes. We study multiparty turn-taking on the VoxConverse dataset and propose an audio-only two-stage pipeline that separates when to trigger a turn boundary from whether the floor is actually transferring. A fast trigger scans the audio and proposes candidate end-of-turn times, while a lightweight verifier runs only at those times to decide \textsc{Hold} or \textsc{Shift} and support next-speaker prediction. We report results in the full multiparty setting and a controlled dyadic top-2 projection for comparability. We also investigate diffusion-based, label-preserving background-audio mixing as a data augmentation strategy. Results show improved shift detection over a baseline, with further improvements from diffusion augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。