arXiv:2509.04147cs.SD2025-09被引 1

用图神经网络优化语音聚类,提升自监督说话人验证效果

Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN

  • 构建相似性连接图,用GCN挖掘语音样本间关系
  • 在VoxCeleb1数据集上使EER降低至2.34%,性能显著提升
  • 适合研究自监督学习与语音识别的开发者参考

随着语音识别技术的发展,说话人验证(SV)成为身份认证的重要手段。传统方法依赖人工特征提取,而深度学习虽显著提升性能,但标注数据稀缺仍限制其应用。自监督学习通过挖掘大规模无标签数据中的隐含信息,增强模型泛化能力,是关键解决方案。DINO是一种高效自监督方法,通过聚类生成无标签语音数据的伪标签以支持后续训练。然而,聚类可能产生噪声伪标签,影响整体识别性能。为此,本文提出基于相似性连接图与图卷积网络(GCN)的改进聚类框架。利用GCN建模结构化数据,并融合相似性连接图中节点间的关联信息,优化聚类过程,提升伪标签准确性,增强自监督说话人验证系统的鲁棒性与性能。实验结果表明,该方法显著提升系统表现,在VoxCeleb1数据集上达到2.34%的等错误率(EER),为自监督说话人验证提供了新思路。

原文摘要 · Abstract (English)

With the continuous development of speech recognition technology, speaker verification (SV) has become an important method for identity authentication. Traditional SV methods rely on handcrafted feature extraction, while deep learning has significantly improved system performance. However, the scarcity of labeled data still limits the widespread application of deep learning in SV. Self-supervised learning, by mining latent information in large unlabeled datasets, enhances model generalization and is a key technology to address this issue. DINO is an efficient self-supervised learning method that generates pseudo-labels from unlabeled speech data through clustering, supporting subsequent training. However, clustering may produce noisy pseudo-labels, which can reduce overall recognition performance. To address this issue, this paper proposes an improved clustering framework based on similarity connection graphs and Graph Convolutional Networks. By leveraging GCNs' ability to model structured data and incorporating relational information between nodes in the similarity connection graph, the clustering process is optimized, improving pseudo-label accuracy and enhancing the robustness and performance of the self-supervised speaker verification system. Experimental results show that this method significantly improves system performance and provides a new approach for self-supervised speaker verification. Index Terms: Speaker Verification, Self-Supervised Learning, DINO, Clustering Algorithm, Graph Convolutional Network, Similarity Connection Graph

说话人验证自监督学习图神经网络语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。