arXiv:2410.01469cs.SDcs.AI2024-10ICLR被引 20

TIGER模型大幅降低语音分离的参数与计算量,同时保持领先性能。

TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation

论文配图:TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
图 1 · 摘自论文原文
  • 通过时频交错增益提取与重建,压缩频域信息并减少计算开销。
  • 参数减少94.3%、MACs降低95.3%,在真实数据上超越SOTA模型TF-GridNet。
  • 引入新数据集EchoSet,提升模型在复杂声学环境下的泛化能力,适合低延迟场景应用。

近年来,语音分离研究多聚焦于提升模型性能,但在低延迟语音处理系统中,高效率同样关键。为此,本文提出一种参数与计算成本显著降低的语音分离模型:时频交错增益提取与重建网络(TIGER)。TIGER利用先验知识对频带进行划分并压缩频域信息,采用多尺度选择性注意力模块提取上下文特征,并引入全频帧注意力模块以捕捉时序与频域上下文信息。此外,为更真实评估模型在复杂声学环境中的表现,本文构建了新数据集EchoSet,包含噪声和更真实的混响(如考虑物体遮挡与材料属性),双说话人语音以随机比例重叠。实验表明,基于EchoSet训练的模型在真实世界数据上具有更强泛化能力,验证了该数据集的实际价值。在EchoSet与真实数据上,TIGER将参数量减少94.3%、MACs降低95.3%,同时性能超越当前最优模型TF-GridNet。

原文摘要 · Abstract (English)

In recent years, much speech separation research has focused primarily on improving model performance. However, for low-latency speech processing systems, high efficiency is equally important. Therefore, we propose a speech separation model with significantly reduced parameters and computational costs: Time-frequency Interleaved Gain Extraction and Reconstruction network (TIGER). TIGER leverages prior knowledge to divide frequency bands and compresses frequency information. We employ a multi-scale selective attention module to extract contextual features while introducing a full-frequency-frame attention module to capture both temporal and frequency contextual information. Additionally, to more realistically evaluate the performance of speech separation models in complex acoustic environments, we introduce a dataset called EchoSet. This dataset includes noise and more realistic reverberation (e.g., considering object occlusions and material properties), with speech from two speakers overlapping at random proportions. Experimental results showed that models trained on EchoSet had better generalization ability than those trained on other datasets compared to the data collected in the physical world, which validated the practical value of the EchoSet. On EchoSet and real-world data, TIGER significantly reduces the number of parameters by 94.3% and the MACs by 95.3% while achieving performance surpassing the state-of-the-art (SOTA) model TF-GridNet.

语音分离高效模型低延迟音频数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。