用声音球谐编码让语音增强模型无视麦克风布局变化。
AmbiDrop: Array-Agnostic Speech Enhancement Using Ambisonics Encoding and Dropout-Based Learning
- 将任意麦克风阵列信号转为球谐域表示,实现几何无关处理。
- 在未见过的阵列布局上仍保持性能,SI-SDR提升1.2dB,PESQ提升0.15。
- 无需多布局数据集,通过通道丢弃提升鲁棒性,适合实际部署。
多通道语音增强利用空间线索提升可懂度与质量,但现有学习方法依赖特定麦克风阵列几何结构,难以适应布局变化。当前无几何依赖方法虽采用大规模多布局数据集,仍可能无法泛化至未见布局。本文提出AmbiDrop(基于丢弃的声学全向编码),通过声学全向信号匹配(ASM)将任意阵列录音编码至球谐域,训练时结合通道丢弃策略以增强对阵列相关编码误差的鲁棒性,从而无需依赖多样化的麦克风阵列数据库。实验表明,基线模型在训练阵列上表现良好,但在未见阵列上性能下降;而AmbiDrop在未见阵列上持续提升SI-SDR、PESQ与STOI指标,展现出强泛化能力与实际应用潜力。
原文摘要 · Abstract (English)
Multichannel speech enhancement leverages spatial cues to improve intelligibility and quality, but most learning-based methods rely on specific microphone array geometry, unable to account for geometry changes. To mitigate this limitation, current array-agnostic approaches employ large multi-geometry datasets but may still fail to generalize to unseen layouts. We propose AmbiDrop (Ambisonics with Dropouts), an Ambisonics-based framework that encodes arbitrary array recordings into the spherical harmonics domain using Ambisonics Signal Matching (ASM). A deep neural network is trained on simulated Ambisonics data, combined with channel dropout for robustness against array-dependent encoding errors, therefore omitting the need for a diverse microphone array database. Experiments show that while the baseline and proposed models perform similarly on the training arrays, the baseline degrades on unseen arrays. In contrast, AmbiDrop consistently improves SI-SDR, PESQ, and STOI, demonstrating strong generalization and practical potential for array-agnostic speech enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。