利用重叠帧信息融合与因果自注意力提升语音增强效果
Speech Enhancement with Overlapped-Frame Information Fusion and Causal Self-Attention
- 通过构造伪重叠帧融合原帧,缓解时频域语音增强的延迟问题
- 引入因果时频通道注意力块,增强模型对多维度特征的表征能力
- 适合需要低延迟高保真语音增强的应用场景
在时频域语音增强方法中,逆时频变换中的重叠相加操作不可避免地引入等于窗长的算法延迟。然而,典型的因果语音增强系统未能利用这一固有延迟内的未来语音信息,从而限制了性能。本文提出一种重叠帧信息融合方案:在每个帧索引处,构造多个伪重叠帧,将其与原始语音帧融合后输入语音增强模型。此外,引入因果时频通道注意力(TFCA)模块,通过在时间、频率和通道维度上并行执行基于自注意力的操作,提升神经网络的表示能力。实验表明,这些改进显著提升了性能,所提出的语音增强系统优于现有先进方法。
原文摘要 · Abstract (English)
For time-frequency (TF) domain speech enhancement (SE) methods, the overlap-and-add operation in the inverse TF transformation inevitably leads to an algorithmic delay equal to the window size. However, typical causal SE systems fail to utilize the future speech information within this inherent delay, thereby limiting SE performance. In this paper, we propose an overlapped-frame information fusion scheme. At each frame index, we construct several pseudo overlapped-frames, fuse them with the original speech frame, and then send the fused results to the SE model. Additionally, we introduce a causal time-frequency-channel attention (TFCA) block to boost the representation capability of the neural network. This block parallelly processes the intermediate feature maps through self-attention-based operations in the time, frequency, and channel dimensions. Experiments demonstrate the superiority of these improvements, and the proposed SE system outperforms the current advanced methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。