arXiv:2607.18658eess.AScs.SD2026-07中稿 · Interspeech 2026

用麦克风位置信息让语音增强模型适配不同麦克风布局

Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

  • 根据麦克风坐标动态调整卷积,显式利用阵列几何信息
  • 在多种麦克风布局下,性能稳定优于传统固定阵列模型
  • 适合需要跨设备部署的语音增强系统研发者

多通道语音增强系统性能优于单通道方法,但受限于固定的麦克风阵列配置,难以在不同设备间部署。尽管近年已有阵列无关的语音增强方法处理麦克风数量和排列的变化,但大多未能利用可用的显式阵列几何先验,错失了最优空间滤波的关键线索。本文提出一种几何感知动态卷积(Geo-DConv)框架,显式利用麦克风坐标,将标准固定阵列语音增强模型转化为鲁棒的阵列不变系统。在最新真实录制的RealMAN多通道语音数据集上进行实验,结果表明,所提架构使两种广泛应用的固定阵列模型能够适应阵列不变设置,在多种阵列拓扑结构下均实现一致的性能提升。

原文摘要 · Abstract (English)

Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they largely fail to exploit explicit array geometry priors when available, missing a crucial cue for optimal spatial filtering. A Geometry-Aware Dynamic Convolution (Geo-DConv) framework is proposed, which explicitly leverages microphone coordinates to transform standard fixed-array SE models into robust array-invariant systems. Experiments are conducted on the recent real-recorded RealMAN multi-channel speech dataset. Results demonstrate that the proposed architecture enables two widely used fixed-array models to adapt to array-invariant settings, with consistent performance improvements across diverse array topologies.

语音增强动态卷积阵列不变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。