arXiv:2607.02296eess.AS2026-07综述

系统梳理声源定位、定向增强与语音识别的融合方法,助力复杂环境下的语音理解。

Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition

论文配图:Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition
图 1 · 摘自论文原文
  • 融合麦克风阵列与学习模型,实现声源定位与方向增强
  • 在混响和噪声下提升语音识别准确率,支持多人场景处理
  • 适用于助听器、智能音箱等实时语音系统,适合开发者参考

真实声学环境下实现鲁棒语音理解仍是机器人听觉、助听器、视频会议系统、智能音箱及语音助手等智能听觉系统的核心挑战。这些系统需应对背景噪声、混响、多说话人干扰和动态声学条件。空间语音感知通过利用麦克风阵列信息,在复杂声场中实现目标语音的定位、增强与解析。本文综述空间语音感知系统,重点关注声源定位(SSL)、定向语音增强(DSE)与自动语音识别(ASR)的作用,涵盖其独立应用及集成处理流程。回顾经典信号处理方法与近年基于学习的方法,包括麦克风阵列定位、波束成形、神经增强、语音分离及现代识别架构。除组件级分析外,还讨论了对噪声与混响的鲁棒性、多说话人处理能力、实时性约束与计算效率。最后分析机器人听觉、听力辅助、智能音箱与视频会议中的典型应用,指出当前开放挑战与未来方向:构建更鲁棒、低延迟、感知驱动的语音系统以应对复杂声学环境。

原文摘要 · Abstract (English)

Robust speech understanding in real-world acoustic environments remains a fundamental challenge for intelligent auditory systems such as robot audition, hearing aids, teleconferencing systems, smart speakers, and voice-controlled assistants. These systems must operate under background noise, reverberation, competing speakers, and dynamic acoustic conditions. Spatial speech perception addresses this challenge by exploiting microphone-array information to localize, enhance, and interpret target speech in complex acoustic scenes. This paper surveys spatial speech perception systems with emphasis on the roles of sound source localization (SSL), directional speech enhancement (DSE), and automatic speech recognition (ASR), both individually and within integrated processing pipelines. We review classical signal-processing approaches and recent learning-based methods for microphone-array localization, beamforming, neural enhancement, speech separation, and modern recognition architectures. Beyond component-level analysis, we discuss robustness to noise and reverberation, multi-speaker operation, real-time constraints, and computational efficiency. We also examine representative applications in robot audition, hearing assistance, smart speakers, and teleconferencing, and identify open challenges and future directions toward robust, low-latency, and perception-aware speech systems for complex acoustic environments.

语音识别声源定位智能音箱助听器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。