arXiv:2509.20410eess.AScs.SD2025-09被引 3

用大模型实现流式语义语音结束点检测,让对话更自然。

Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction

  • 基于大模型语义理解与滑动窗口训练,支持流式推理。
  • 在完整与不完整语音场景下表现优秀,效果领先。
  • 可独立优化,适合下一代全双工人机交互系统。

语音对话模型显著提升了人机交互智能化水平,但缺乏即插即用的全双工语义结束点检测模块,制约了音频交互的流畅性。本文提出Phoenix-VAD,一种基于大语言模型(LLM)的流式语义结束点检测方法。该方法利用大模型的语义理解能力与滑动窗口训练策略,在支持流式推理的同时实现可靠的语义结束点检测。在语义完整与不完整语音场景下的实验表明,Phoenix-VAD表现优异且具有竞争力。此外,该设计使全双工预测模块可独立于对话模型进行优化,为下一代人机交互提供更可靠、灵活的支持。

原文摘要 · Abstract (English)

Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detection, hindering seamless audio interactions. In this paper, we introduce Phoenix-VAD, an LLM-based model that enables streaming semantic endpoint detection. Specifically, Phoenix-VAD leverages the semantic comprehension capability of the LLM and a sliding window training strategy to achieve reliable semantic endpoint detection while supporting streaming inference. Experiments on both semantically complete and incomplete speech scenarios indicate that Phoenix-VAD achieves excellent and competitive performance. Furthermore, this design enables the full-duplex prediction module to be optimized independently of the dialogue model, providing more reliable and flexible support for next-generation human-computer interaction.

语音交互语义检测大模型流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。