轻量级语音增强网络,适配任意麦克风阵列,实时高效。
LABNet: A Lightweight Attentive Beamforming Network for Ad-hoc Multichannel Microphone Invariant Real-Time Speech Enhancement
- 三阶段架构,跨通道注意力融合多麦克风特征
- 仅需1.2%计算量,实现麦克风无关性与高保真增强
- 适合边缘设备部署,尤其适用于移动/可穿戴场景
多通道语音增强旨在利用时空信号特征从噪声中恢复清晰语音。在即兴麦克风阵列条件下,麦克风无关性(MI)要求系统能处理不同数量和布局的麦克风。然而,实际应用中多通道录音会显著增加边缘设备的计算负担,因此亟需轻量化、高效的部署方案。本文提出一种轻量级注意力波束成形网络(LABNet),在低复杂度实时语音增强系统中集成麦克风无关性。设计了三阶段框架,实现高效的通道内建模与通道间交互;引入跨通道注意力模块,选择性聚合各通道特征。实验表明,该方法在保持麦克风无关性的同时,仅需1.2%的计算开销,性能优异,展现出在即兴阵列处理中的巨大潜力。代码已开源:https://github.com/Jokejiangv/LABNet.git
原文摘要 · Abstract (English)
Multichannel speech enhancement (SE) aims to restore clean speech from noisy measurements by leveraging spatiotemporal signal features. In ad-hoc array conditions, microphone invariance (MI) requires systems to handle different microphone numbers and array geometries. From a practical perspective, multichannel recordings inevitably increase the computational burden for edge-device applications, highlighting the necessity of lightweight and efficient deployments. In this work, we propose a lightweight attentive beamforming network (LABNet) to integrate MI in a low-complexity real-time SE system. We design a three-stage framework for efficient intra-channel modeling and inter-channel interaction. A cross-channel attention module is developed to aggregate features from each channel selectively. Experimental results demonstrate our LABNet achieves impressive performance with ultra-light resource overhead while maintaining the MI, indicating great potential for ad-hoc array processing. The code is available:https://github.com/Jokejiangv/LABNet.git
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。