arXiv:2508.17336cs.SDcs.AI2025-08被引 1

融合体声与空气麦克风,动态优化降噪与高频还原

Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework

  • 用映射与掩码网络分别增强体声和空气麦克风信号
  • 在TAPS+DNS-2023数据集上,多种噪声下性能超越单模态
  • 适合需要高鲁棒性语音采集的可穿戴设备场景

体声麦克风信号(BMS)绕过空气传播,具备强抗噪能力,但存在高频信息丢失问题。本研究提出一种新型多模态框架,结合体声麦克风(BMS)与空气麦克风信号(AMS),实现降噪与高频重建。不同于传统特征拼接方法,该框架采用两个专用网络:基于映射的模型用于增强BMS,基于掩码的模型用于去噪AMS。通过动态融合机制根据局部噪声条件自适应选择最优模态贡献,充分发挥各自优势。在包含DNS-2023噪声片段的TAPS数据集上进行评估,使用客观语音质量指标测试。结果表明,该方法在多种噪声环境下均显著优于单一模态方案。

原文摘要 · Abstract (English)

Body-conduction microphone signals (BMS) bypass airborne sound, providing strong noise resistance. However, a complementary modality is required to compensate for the inherent loss of high-frequency information. In this study, we propose a novel multi-modal framework that combines BMS and acoustic microphone signals (AMS) to achieve both noise suppression and high-frequency reconstruction. Unlike conventional multi-modal approaches that simply merge features, our method employs two specialized networks: a mapping-based model to enhance BMS and a masking-based model to denoise AMS. These networks are integrated through a dynamic fusion mechanism that adapts to local noise conditions, ensuring the optimal use of each modality's strengths. We performed evaluations on the TAPS dataset, augmented with DNS-2023 noise clips, using objective speech quality metrics. The results clearly demonstrate that our approach outperforms single-modal solutions in a wide range of noisy environments.

语音增强多模态体声麦克风

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。