arXiv:2607.00548eess.AS2026-07

用声音空间编码技术让语音增强模型无视麦克风布局,适合各种设备部署。

AmbiDrop: Ambisonics-Based Array-Agnostic Neural Speech Enhancement

论文配图:AmbiDrop: Ambisonics-Based Array-Agnostic Neural Speech Enhancement
图 1 · 摘自论文原文
  • 用通道丢弃模拟声场编码误差,让模型学会忽略麦克风物理位置。
  • 在多种未见过的麦克风布局上表现稳定,真实录音测试准确率超90%。
  • 模型轻量且抗传感器故障,适合低功耗可穿戴设备使用。

多通道深度神经网络显著提升了语音增强性能,但通常依赖固定的麦克风阵列几何结构,导致在未见或不规则配置下泛化能力差。现有无布局依赖方法常需高复杂度架构或海量多样数据集,仍难以泛化到分布外布局。本文深入分析AmbiDrop框架,该框架通过将理想声场编码(Ambisonics)作为DNN输入实现几何无关性。训练时采用通道级丢弃层模拟声场编码误差,使学习过程与物理传感器排列解耦。推理阶段,任意阵列配置的麦克风信号通过声场信号匹配(ASM)转换至声场域进行处理。大量实验表明,AmbiDrop在多种未见过的仿真阵列和真实录音中保持高鲁棒性。此外,结果表明该框架对传感器失效具有韧性,即使在网络规模缩减时仍有效,非常适用于资源受限的边缘设备和多功能可穿戴硬件部署。

原文摘要 · Abstract (English)

Multichannel Deep Neural Networks (DNNs) have significantly improved speech enhancement performance; however, they typically remain constrained by reliance on fixed microphone array geometries, leading to poor generalization on unseen or irregular configurations. Current array-agnostic approaches often rely on high-complexity architectures or massive, diverse datasets, yet they still struggle to generalize to out-of-distribution layouts. In this paper, we present an in-depth analysis of AmbiDrop, a recently proposed framework that achieves geometry independence by leveraging ideal Ambisonics as the DNN input. By employing a channel-wise dropout layer during training to simulate Ambisonics encoding errors, AmbiDrop decouples the learning process from the physical sensor arrangement. During inference, microphone signals from arbitrary array configurations are transformed into the Ambisonics domain via Ambisonics Signal Matching (ASM) before processing. Extensive experiments demonstrate that AmbiDrop maintains high robustness across a diverse suite of unseen simulated arrays and real-world recordings. Furthermore, our results show that the framework is resilient to sensor failures and remains effective even with reduced network scales, making it highly suitable for deployment on resource-constrained edge devices and versatile wearable hardware.

语音增强声场编码边缘计算可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。