arXiv:2603.26515cs.CLcs.AI2026-03被引 1

用声学与语言联合建模,实现低延迟高精度对话轮换检测。

JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems

  • 融合声学与语言特征,通过交叉注意力动态整合表示。
  • 在多语种公开数据集和日语客服数据上准确率领先现有方法。
  • 可与语音识别并行运行,无额外延迟,适合工业级实时系统。

尽管取得进展,工业级语音AI部署中高效且鲁棒的对话轮换检测仍是重大挑战。现有系统多仅依赖声学或语义线索,导致准确性与稳定性不足;近期尝试赋予大语言模型全双工能力需大量全双工数据,训练与部署开销大,影响实时性。本文提出JAL-Turn,一种轻量高效的纯语音对话轮换检测框架,采用声学-语言联合建模范式,通过交叉注意力模块自适应融合预训练声学表征与语言特征,支持低延迟的保持(hold)与切换(shift)状态预测。通过共享冻结的ASR编码器,JAL-Turn使轮换检测可完全与语音识别并行执行,不引入额外端到端延迟或计算开销。此外,我们构建了可扩展的数据生成管道,从大规模真实对话语料中自动标注可靠轮换标签。在多个公开多语种基准及内部日语客服数据集上的大量实验表明,JAL-Turn在检测准确率上持续优于强基线,同时保持卓越的实时性能。

原文摘要 · Abstract (English)

Despite recent advances, efficient and robust turn-taking detection remains a significant challenge in industrial-grade Voice AI agent deployments. Many existing systems rely solely on acoustic or semantic cues, leading to suboptimal accuracy and stability, while recent attempts to endow large language models with full-duplex capabilities require costly full-duplex data and incur substantial training and deployment overheads, limiting real-time performance. In this paper, we propose JAL-Turn, a lightweight and efficient speech-only turn-taking framework that adopts a joint acoustic-linguistic modeling paradigm, in which a cross-attention module adaptively integrates pre-trained acoustic representations with linguistic features to support low-latency prediction of hold vs shift states. By sharing a frozen ASR encoder, JAL-Turn enables turn-taking prediction to run fully in parallel with speech recognition, introducing no additional end-to-end latency or computational overhead. In addition, we introduce a scalable data construction pipeline that automatically derives reliable turn-taking labels from large-scale real-world dialogue corpora. Extensive experiments on public multilingual benchmarks and an in-house Japanese customer-service dataset show that JAL-Turn consistently outperforms strong state-of-the-art baselines in detection accuracy while maintaining superior real-time performance.

对话系统语音识别实时检测联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。