改进TDNN的上下文建模能力,显著降低语音验证错误率
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
- 设计新结构增强多尺度上下文特征提取
- 在VoxCeleb1-O上错误率降低近23%
- 适合关注语音验证模型优化的研究者
目前,时延神经网络(TDNN)已成为语音验证任务的主流架构,其中ECAPA-TDNN是性能领先的模型。现有研究主要聚焦于提升TDNN对全局信息的建模能力,并缩小其与二维卷积的差距。然而,ECAPA-TDNN中采用的SE-Res2Block层级卷积结构未能充分挖掘上下文信息,导致其建模有效上下文依赖的能力较弱。为此,本文提出三种基于ECAPA-TDNN的改进架构,旨在更充分、有效地提取具有上下文依赖性的多尺度特征,并进行特征聚合。在VoxCeleb和CN-Celeb数据集上的实验结果验证了所提方法的有效性。其中一种架构在VoxCeleb1-O数据集上的等错误率(EER)相比ECAPA-TDNN降低了近23%,在参数量相当的前提下展现出当前TDNN架构中的竞争性性能。
原文摘要 · Abstract (English)
Today, Time Delay Neural Network (TDNN) has become the mainstream architecture for speaker verification task, in which the ECAPA-TDNN is one of the state-of-the-art models. The current works that focus on improving TDNN primarily address the limitations of TDNN in modeling global information and bridge the gap between TDNN and 2-Dimensional convolutions. However, the hierarchical convolutional structure in the SE-Res2Block proposed by ECAPA-TDNN cannot make full use of the contextual information, resulting in the weak ability of ECAPA-TDNN to model effective context dependencies. To this end, three improved architectures based on ECAPA-TDNN are proposed to fully and effectively extract multi-scale features with context dependence and then aggregate these features. The experimental results on VoxCeleb and CN-Celeb verify the effectiveness of the three proposed architectures. One of these architectures achieves nearly a 23% lower Equal Error Rate compared to that of ECAPA-TDNN on VoxCeleb1-O dataset, demonstrating the competitive performance achievable among the current TDNN architectures under the comparable parameter count.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。