arXiv:2604.01897cs.SDeess.AS2026-04被引 3

融合声学与语义线索,实现低延迟高鲁棒的对话轮次检测。

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection

论文配图:FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
图 1 · 摘自论文原文
  • 结合流式CTC与声学特征,从部分输入早期做出判断。
  • 在真实对话数据上准确率更高,中断延迟更低。
  • 适合需要实时交互的语音助手、客服系统等场景。

近期音频大模型的发展使语音对话系统从基于轮次的交互迈向实时全双工通信,要求智能体在用户说话时即时判断何时发言、让位或打断。现有全双工方法要么依赖仅含声学信息的语音活动检测(缺乏语义理解),要么依赖自动语音识别(ASR)模块,带来延迟并在重叠语音和噪声环境下性能下降。此外,现有数据集很少捕捉真实的交互动态,限制了评估与部署。为此,我们提出 extbf{FastTurn},一种统一的低延迟、高鲁棒的轮次检测框架。为降低延迟并保持性能,FastTurn 结合流式CTC解码与声学特征,实现基于部分观测的早期决策,同时保留语义线索。我们还发布了一个基于真实人类对话构建的测试集,包含真实的轮次转换、重叠语音、回应词、停顿、音调变化和环境噪声。实验表明,FastTurn 在代表性基线之上实现了更高的决策准确率和更低的中断延迟,并在复杂声学条件下保持鲁棒性,证明其在实际全双工对话系统中的有效性。

原文摘要 · Abstract (English)

Recent advances in AudioLLMs have enabled spoken dialogue systems to move beyond turn-based interaction toward real-time full-duplex communication, where the agent must decide when to speak, yield, or interrupt while the user is still talking. Existing full-duplex approaches either rely on voice activity cues, which lack semantic understanding, or on ASR-based modules, which introduce latency and degrade under overlapping speech and noise. Moreover, available datasets rarely capture realistic interaction dynamics, limiting evaluation and deployment. To mitigate the problem, we propose \textbf{FastTurn}, a unified framework for low-latency and robust turn detection. To advance latency while maintaining performance, FastTurn combines streaming CTC decoding with acoustic features, enabling early decisions from partial observations while preserving semantic cues. We also release a test set based on real human dialogue, capturing authentic turn transitions, overlapping speech, backchannels, pauses, pitch variation, and environmental noise. Experiments show FastTurn achieves higher decision accuracy with lower interruption latency than representative baselines and remains robust under challenging acoustic conditions, demonstrating its effectiveness for practical full-duplex dialogue systems.

对话系统低延迟语音识别全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。