arXiv:2409.11725eess.AScs.SD2024-09被引 1

超轻量语音增强模型,14K参数适配边缘设备

Dense-TSNet: Dense Connected Two-Stage Structure for Ultra-Lightweight Speech Enhancement

  • 用密集连接两阶段结构提升训练稳定性
  • 仅14K参数实现良好降噪效果
  • 适合低算力设备部署,如手机端

语音增强旨在改善噪声环境下的语音质量与可懂度。近期研究多采用深度神经网络,特别是两阶段(TS)架构以增强特征提取,但模型复杂度与规模仍较大,限制了在资源受限场景的应用。为解决此问题,本文提出Dense-TSNet,一种新型超轻量语音增强网络。该方法采用新颖的密集两阶段(Dense-TS)架构,在后期训练中更稳健地优化目标函数,克服基线模型过早收敛的问题。同时引入多视角凝视块(MVGB),通过卷积网络融合全局、通道与局部信息以增强特征提取。此外,分析了损失函数对感知质量的影响。Dense-TSNet仅需约14K参数,表现出优异性能,特别适合部署于资源受限环境。

原文摘要 · Abstract (English)

Speech enhancement aims to improve speech quality and intelligibility in noisy environments. Recent advancements have concentrated on deep neural networks, particularly employing the Two-Stage (TS) architecture to enhance feature extraction. However, the complexity and size of these models remain significant, which limits their applicability in resource-constrained scenarios. Designing models suitable for edge devices presents its own set of challenges. Narrow lightweight models often encounter performance bottlenecks due to uneven loss landscapes. Additionally, advanced operators such as Transformers or Mamba may lack the practical adaptability and efficiency that convolutional neural networks (CNNs) offer in real-world deployments. To address these challenges, we propose Dense-TSNet, an innovative ultra-lightweight speech enhancement network. Our approach employs a novel Dense Two-Stage (Dense-TS) architecture, which, compared to the classic Two-Stage architecture, ensures more robust refinement of the objective function in the later training stages. This leads to improved final performance, addressing the early convergence limitations of the baseline model. We also introduce the Multi-View Gaze Block (MVGB), which enhances feature extraction by incorporating global, channel, and local perspectives through convolutional neural networks (CNNs). Furthermore, we discuss how the choice of loss function impacts perceptual quality. Dense-TSNet demonstrates promising performance with a compact model size of around 14K parameters, making it particularly well-suited for deployment in resource-constrained environments.

语音增强轻量化边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。