用说话间隔时间预测提升流式语音端点检测准确率
Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction

- 以说话者下一次开口时间作为训练目标,无需额外标注
- 320毫秒内端点检测准确率比最强基线提升25.9%
- 对长停顿场景效果更优,适合实时语音交互系统
端点检测(EPD)对于实现流式语音系统中的自然对话轮换至关重要。然而,由于说话人常因犹豫或不流畅而中途停顿,可靠判断语句终点仍具挑战性。语义端点检测虽有前景,但受限于模糊的监督信号和严格的流式处理约束。本文提出Next-Turn方法,以「到下一次说话开始的时间」为训练目标,目标直接从语音时间戳中获取,无需额外标注。实验表明,该方法在320毫秒内相比最强基线,端点检测准确率提升25.9%。此外,与持续时间感知目标联合训练可补充传统二分类端点检测,在停顿时间越长时性能增益越显著。
原文摘要 · Abstract (English)
Endpoint detection (EPD) is essential for natural turn-taking in streaming speech systems. However, reliably determining the endpoint of an utterance is challenging because speakers often pause mid-utterance due to hesitations and disfluencies. Semantic EPD has emerged as a promising direction to address this issue but is hindered by ambiguous supervision and strict streaming constraints. We propose Next-Turn that uses the time-to-next-speech-onset as the training objective, where targets are derived directly from speech timestamps and require no additional annotation. Experiments show that the proposed method outperforms conventional acoustic and recent semantic EPD baselines, achieving a 25.9% absolute improvement in endpoint accuracy within 320 ms over the strongest baseline. In addition, joint training with the duration-aware objective complements standard binary EPD, with gains that increase monotonically with increasing pauses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。