用低比特神经音频编码器实现低延迟对话端点检测
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
- 用神经音频编码器提取流式音频特征,支持实时处理
- 在160毫秒延迟下,误切率降低超37%
- 适合需要快速响应的语音对话系统开发者
准确且低延迟的端点检测对有效语音对话系统至关重要。传统方法多依赖频谱特征,本文提出基于流式、低比特神经音频编码器(NAC)特征的多轮对话实时语音端点检测方法,利用神经音频编码器最新进展。为减少误切错误,引入新型标签延迟训练策略。在固定中位延迟160毫秒条件下,结合NAC与标签延迟的方法相比基线模型,单流端点检测器误切率降低42.7%,双流配置降低37.5%。最后,该方法成功集成至基于编码器的预训练语音大语言模型,使其平均响应时间缩短1200毫秒,误切率下降35%。
原文摘要 · Abstract (English)
Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。