通过多粒度融合提升说话人验证的细粒度特征捕捉能力
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
- 采用深度可分离卷积与多粒度融合结构,兼顾时频域细节与全局上下文
- 在VoxCeleb数据集上达到顶尖性能,参数量和计算量均保持高效
- 适合需要高精度且资源受限的实时说话人验证场景
在说话人验证中,传统模型侧重建模长时上下文特征以捕捉全局说话人特性,但常忽略包含高度判别性信息的细粒度语音指纹。本文提出一种新型模型架构MGFF-TDNN,基于多粒度特征融合。该模型使用二维深度可分离卷积模块作为前端特征提取器,增强局部特征建模能力,有效捕捉时频域特征。为实现全面的多粒度特征融合,提出M-TDNN结构,通过结合时延神经网络与音素级特征池化,同时实现全局上下文建模与细粒度特征提取。在VoxCeleb数据集上的实验表明,MGFF-TDNN在说话人验证任务中表现优异,且参数量与计算资源消耗保持高效。
原文摘要 · Abstract (English)
In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly discriminative features essential for robust speaker embeddings. This paper introduces a novel model architecture, termed MGFF-TDNN, based on multi-granularity feature fusion. The MGFF-TDNN leverages a two-dimensional depth-wise separable convolution module, enhanced with local feature modeling, as a front-end feature extractor to effectively capture time-frequency domain features. To achieve comprehensive multi-granularity feature fusion, we propose the M-TDNN structure, which integrates global contextual modeling with fine-grained feature extraction by combining time-delay neural networks and phoneme-level feature pooling. Experiments on the VoxCeleb dataset demonstrate that the MGFF-TDNN achieves outstanding performance in speaker verification while remaining efficient in terms of parameters and computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。