用空间滤波器组提升麦克风阵列语音增强的泛化能力
Spatial-Filter-Bank-Based Neural Method for Multichannel Speech Enhancement
- 基于空间滤波器组提取对几何参数不敏感的特征
- 在固定阵列上训练,跨未见阵列配置仍保持效果
- 适合实际部署中阵列位置不确定的场景
基于深度学习的多通道语音增强方法在麦克风阵列几何参数变化时性能往往下降。传统方法通常需在多个阵列上训练,成本较高。本文聚焦均匀圆形阵列,提出使用空间滤波器组提取近似几何参数不变的特征,并通过两阶段Conformer模型(TSCBM)进行语音增强。实验表明,该方法仅在固定阵列上训练,即可在应用时有效应对未见过的均匀圆形阵列几何配置。
原文摘要 · Abstract (English)
The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on multiple microphone arrays, which can be costly. To address this challenge, we focus on uniform circular arrays and propose the use of a spatial filter bank to extract features that are approximately invariant to geometric parameters. These features are then processed by a two-stage conformer-based model (TSCBM) to enhance speech quality. Experimental results demonstrate that our proposed method can be trained on a fixed microphone array while maintaining effective performance across uniform circular arrays with unseen geometric configurations during applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。