arXiv:2510.15659eess.AS2025-10

用协同注意力融合相位与幅度特征,提升说话人识别准确率

Magnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognition

  • 通过双分支网络分别处理幅度和相位特征
  • 协同注意力动态调整两域权重,提升识别精度至97.20%
  • 适合需要高精度说话人验证的语音系统开发

基于相位的声源特征可融入基于幅度的说话人识别系统以提升性能。然而,传统特征级融合方法常忽略幅度与相位域中说话人语义的独特贡献。为此,本文提出一种基于协同注意力机制的特征级融合框架。该框架包含两个独立子网络,分别处理幅度与相位域;随后在池化层前,利用协同注意力机制融合两域的高层中间输出。协同注意力模块生成的相关矩阵可依据不同发音动态重分配幅度与相位域的权重。在VoxCeleb数据集上的实验表明,该融合策略将Top-1准确率提升至97.20%,相比当前最优系统绝对提升0.82%,相较仅使用FBank的单特征系统,EER降低0.45%。

原文摘要 · Abstract (English)

Phase-based features related to vocal source characteristics can be incorporated into magnitude-based speaker recognition systems to improve the system performance. However, traditional feature-level fusion methods typically ignore the unique contributions of speaker semantics in the magnitude and phase domains. To address this issue, this paper proposed a feature-level fusion framework using the co-attention mechanism for speaker recognition. The framework consists of two separate sub-networks for the magnitude and phase domains respectively. Then, the intermediate high-level outputs of both domains are fused by the co-attention mechanism before a pooling layer. A correlation matrix from the co-attention module is supposed to re-assign the weights for dynamically scaling contributions in the magnitude and phase domains according to different pronunciations. Experiments on VoxCeleb showed that the proposed feature-level fusion strategy using the co-attention mechanism gave the Top-1 accuracy of 97.20%, outperforming the state-of-the-art system with 0.82% absolutely, and obtained EER reduction of 0.45% compared to single feature system using FBank.

说话人识别协同注意力特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。