arXiv:2512.22148cs.SDcs.AI2025-12中稿 · Interspeech 2025被引 2

用动态加权的层注意力池化提升语音验证精度

Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification

  • 提出层注意力池化,动态评估各层重要性并用最大池化融合特征
  • 在VoxCeleb上达到顶尖性能,训练时间大幅缩短
  • 适合需要高效高精度语音验证的工程应用

近期语音验证研究通过利用预训练Transformer模型的逐层输出取得了显著进展。然而,关于如何超越静态加权平均来聚合多层级特征的研究仍较少。本文提出层注意力池化(LAP),一种用于语音验证的新型多层表示聚合策略。LAP从多个角度时变地评估各层重要性,并采用最大池化而非平均池化。此外,我们设计了一种轻量级后端说话人模型,包含LAP和注意力统计时序池化(ASTP),用于从预训练模型输出中提取说话人嵌入。在VoxCeleb基准上的实验表明,该紧凑架构在大幅减少训练时间的同时实现了最先进的性能。我们进一步分析了LAP设计及其动态加权机制在捕捉说话人特征方面的有效性。

原文摘要 · Abstract (English)

Recent speaker verification studies have achieved notable success by leveraging layer-wise output from pre-trained Transformer models. However, few have explored the advancements in aggregating these multi-level features beyond the static weighted average. We present Layer Attentive Pooling (LAP), a novel strategy for aggregating inter-layer representations from pre-trained speech models for speaker verification. LAP assesses the significance of each layer from multiple perspectives time-dynamically, and employs max pooling instead of averaging. Additionally, we propose a lightweight backend speaker model comprising LAP and Attentive Statistical Temporal Pooling (ASTP) to extract speaker embeddings from pre-trained model output. Experiments on the VoxCeleb benchmark reveal that our compact architecture achieves state-of-the-art performance while greatly reducing the training time. We further analyzed LAP design and its dynamic weighting mechanism for capturing speaker characteristics.

语音验证注意力机制特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。