利用声音方向信息提升嘈杂环境下的关键词识别准确率
End-to-End Direction-Aware Keyword Spotting with Spatial Priors in Noisy Environments
- 多通道输入结合空间编码与方向先验,端到端优化
- 在不同信噪比下,联合使用空间建模与方向先验效果最佳
- 适合需要精准定位说话人语音的复杂声学场景
关键词识别(KWS)在众多语音驱动应用中至关重要,但在嘈杂环境中的鲁棒性仍具挑战。传统系统通常依赖单通道输入并采用前后端分离的级联架构,难以实现联合优化,性能受限。本文提出一种端到端多通道KWS框架,利用空间线索提升抗噪能力:空间编码器学习通道间特征,空间嵌入注入方向先验,融合表示由流式主干网络处理。在多种信噪比(SNRs)的模拟噪声条件下实验表明,空间建模与方向先验各自带来明显提升,二者结合达到最优效果。结果验证了端到端多通道空间建模的有效性,表明其在复杂声学场景中具备强目标说话人感知检测潜力。
原文摘要 · Abstract (English)
Keyword spotting (KWS) is crucial for many speech-driven applications, but robust KWS in noisy environments remains challenging. Conventional systems often rely on single-channel inputs and a cascaded pipeline separating front-end enhancement from KWS. This precludes joint optimization, inherently limiting performance. We present an end-to-end multi-channel KWS framework that exploits spatial cues to improve noise robustness. A spatial encoder learns inter-channel features, while a spatial embedding injects directional priors; the fused representation is processed by a streaming backbone. Experiments in simulated noisy conditions across multiple signal-to-noise ratios (SNRs) show that spatial modeling and directional priors each yield clear gains over baselines, with their combination achieving the best results. These findings validate end-to-end multi-channel spatial modeling, indicating strong potential for the target-speaker-aware detection in complex acoustic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。