轻量级注意力池化方法,提升语音验证的上下文建模与跨任务泛化能力
CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
- 通过分组可学习查询建模帧间上下文依赖,共享键值保持高效
- 在VoxCeleb上达到0.42%~0.96%的错误率,参数更少、收敛更快
- 适用于多种自监督模型和任务,如情感识别与防欺骗检测
近年来,基于自监督学习(SSL)的说话人验证(SV)系统受到广泛关注。然而,现有方法常难以捕捉局部时间依赖性,并在不同任务间泛化能力不足。本文提出一种轻量级框架——上下文感知多头分解注意力池化(CA-MHFA),通过引入邻近帧的上下文信息来增强建模能力。该方法采用分组可学习查询,有效建模上下文依赖,同时通过组间共享键和值保持计算效率。在VoxCeleb数据集上的实验表明,CA-MHFA在Vox1-O、Vox1-E和Vox1-H上的等错误率(EER)分别为0.42%、0.48%和0.96%,优于如WavLM-TDNN等复杂模型,且参数更少、收敛更快。此外,该方法在多种SSL模型和任务中展现出强泛化能力,包括情感识别与反欺骗任务,体现了其鲁棒性与通用性。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42\%, 0.48\%, and 0.96\% on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。