arXiv:2509.09932eess.AS2025-09被引 1

改进TDNN的上下文建模能力,显著降低语音验证错误率

Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification

  • 设计新结构增强多尺度上下文特征提取
  • 在VoxCeleb1-O上错误率降低近23%
  • 适合关注语音验证模型优化的研究者

目前,时延神经网络(TDNN)已成为语音验证任务的主流架构,其中ECAPA-TDNN是性能领先的模型。现有研究主要聚焦于提升TDNN对全局信息的建模能力,并缩小其与二维卷积的差距。然而,ECAPA-TDNN中采用的SE-Res2Block层级卷积结构未能充分挖掘上下文信息,导致其建模有效上下文依赖的能力较弱。为此,本文提出三种基于ECAPA-TDNN的改进架构,旨在更充分、有效地提取具有上下文依赖性的多尺度特征,并进行特征聚合。在VoxCeleb和CN-Celeb数据集上的实验结果验证了所提方法的有效性。其中一种架构在VoxCeleb1-O数据集上的等错误率(EER)相比ECAPA-TDNN降低了近23%,在参数量相当的前提下展现出当前TDNN架构中的竞争性性能。

原文摘要 · Abstract (English)

Today, Time Delay Neural Network (TDNN) has become the mainstream architecture for speaker verification task, in which the ECAPA-TDNN is one of the state-of-the-art models. The current works that focus on improving TDNN primarily address the limitations of TDNN in modeling global information and bridge the gap between TDNN and 2-Dimensional convolutions. However, the hierarchical convolutional structure in the SE-Res2Block proposed by ECAPA-TDNN cannot make full use of the contextual information, resulting in the weak ability of ECAPA-TDNN to model effective context dependencies. To this end, three improved architectures based on ECAPA-TDNN are proposed to fully and effectively extract multi-scale features with context dependence and then aggregate these features. The experimental results on VoxCeleb and CN-Celeb verify the effectiveness of the three proposed architectures. One of these architectures achieves nearly a 23% lower Equal Error Rate compared to that of ECAPA-TDNN on VoxCeleb1-O dataset, demonstrating the competitive performance achievable among the current TDNN architectures under the comparable parameter count.

语音验证深度学习特征提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。