用新型神经网络提升语音增强效果,参数少一半还更清晰。
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
- 用GR-KAN替代传统层,提升语音信号建模能力
- 参数减少4倍,语音质量评分(PESQ)提升0.1
- 首次在时域和频域均实现性能提升,适合语音处理研究者
基于深度神经网络(DNN)的语音增强(SE)通常使用传统激活函数,难以捕捉高保真增强所需的复杂多尺度结构。群组有理KAN(GR-KAN)作为科尔莫戈罗夫-阿诺尔德网络(KAN)的变体,在保持表达力的同时提升了复杂任务的可扩展性。我们将GR-KAN应用于现有DNN-based SE框架:在时频域的MP-SENet中以GR-KAN层替换全连接层,并将GR-KAN激活函数引入时域的Demucs模型中的1D CNN层。在Voicebank-DEMAND数据集上的实验表明,GR-KAN最多可减少4倍参数量,同时使PESQ评分提升最高0.1。相比之下,原始KAN因可扩展性问题,在小规模信号建模任务中虽优于MLP,但未能改进MP-SENet。本工作首次实现基于KAN的方法在时域与当前最优的频域SE中均取得稳定提升,确立了GR-KAN在语音增强中的潜力。
原文摘要 · Abstract (English)
Deep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures needed for high-fidelity SE. Group-Rational KAN (GR-KAN), a variant of Kolmogorov-Arnold Networks (KAN), retains KAN's expressiveness while improving scalability on complex tasks. We adapt GR-KAN to existing DNN-based SE by replacing dense layers with GR-KAN layers in the time-frequency (T-F) domain MP-SENet and adapting GR-KAN's activations into the 1D CNN layers in the time-domain Demucs. Results on Voicebank-DEMAND show that GR-KAN requires up to 4x fewer parameters while improving PESQ by up to 0.1. In contrast, KAN, facing scalability issues, outperforms MLP on a small-scale signal modeling task but fails to improve MP-SENet. We demonstrate the first successful use of KAN-based methods for consistent improvement in both time- and SoTA TF-domain SE, establishing GR-KAN as a promising alternative for SE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。