arXiv:2501.07215eess.AScs.SD2025-01中稿 · IEEE Signal Proces…被引 16

融合模型与深度学习,提升麦克风阵列语音增强效果

Microphone Array Signal Processing and Deep Learning for Speech Enhancement

  • 结合模型先验与数据驱动,优化空间滤波参数估计
  • 在降噪、分离和去混响任务中表现优于传统方法
  • 适合需要高精度语音增强的智能设备开发者

多通道声学信号处理是利用目标信号与噪声源之间空间差异进行信号增强的有效工具。然而,经典的最优数据依赖空间滤波方法依赖于信号二阶统计矩的知识,这些信息传统上难以获取。本文比较了基于模型、纯数据驱动以及混合方法在参数估计与滤波中的应用,其中混合方法旨在结合模型驱动信号处理与数据驱动深度学习的优势,以克服各自局限。通过降噪、声源分离和去混响等实例,展示了其设计原理。

原文摘要 · Abstract (English)

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal enhancement. However, the textbook solutions for optimal data-dependent spatial filtering rest on the knowledge of second-order statistical moments of the signals, which have traditionally been difficult to acquire. In this contribution, we compare model-based, purely data-driven, and hybrid approaches to parameter estimation and filtering, where the latter tries to combine the benefits of model-based signal processing and data-driven deep learning to overcome their individual deficiencies. We illustrate the underlying design principles with examples from noise reduction, source separation, and dereverberation.

语音增强麦克风阵列深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。