融合骨传导与空气传导语音,提升极低信噪比下的语音增强效果
DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
- 三分支架构通过迭代注意力与门控机制实现跨模态自适应融合
- 在多种噪声下显著提升语音质量与可懂度,下游语音识别错误率降低2.5%以上
- 适合高噪声环境下的实时语音增强,如军事通信或智能助听设备
传统语音增强系统在极低信噪比(SNR)环境下性能急剧下降,因空气传导(AC)麦克风易受环境噪声干扰。尽管骨传导(BC)传感器提供抗噪互补信息,现有融合方法在宽范围SNR条件下难以保持稳定表现。为此,本文提出深度平衡多模态迭代融合框架(DBMIF),采用三分支结构,基于多尺度交互式编码解码骨干网络,引入迭代注意力模块与跨分支门控模块,实现自适应加权与双向信息交换。同时,设计平衡交互瓶颈层,学习紧凑稳定的融合表示。大量实验表明,DBMIF在多种噪声类型下,语音质量与可懂度均优于近期单模态及多模态基线。在下游自动语音识别任务中,字符错误率较现有方法至少降低2.5%。结果证实,DBMIF有效利用了BC语音的鲁棒性,同时保留了AC语音的自然性,确保真实场景下的可靠性。源代码已公开于github.com/wyl516w/dbmif。
原文摘要 · Abstract (English)
The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC) sensors offer complementary, noise-tolerant information, existing fusion approaches struggle to maintain consistent performance across a wide range of SNR conditions. To address this limitation, we propose the Deep Balanced Multimodal Iterative Fusion Framework (DBMIF), a three-branch architecture designed to reconstruct high-fidelity speech through rigorous cross-modal interaction. Specifically, grounded in a multi-scale interactive encoder-decoder backbone, the framework orchestrates an iterative attention module and a cross-branch gated module to facilitate adaptive weighting and bidirectional exchange. To complement this dynamic interaction, a balanced-interaction bottleneck is further integrated to learn a compact, stable fused representation. Extensive experiments demonstrate that DBMIF achieves competitive performance compared with recent unimodal and multimodal baselines in both speech quality and intelligibility across diverse noise types. In downstream ASR tasks, the proposed method reduces the character error rate by at least 2.5 percent compared to competing approaches. These results confirm that DBMIF effectively harnesses the robustness of BC speech while preserving the naturalness of AC speech, ensuring reliability in real-world scenarios. The source code is publicly available at github.com/wyl516w/dbmif.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。