arXiv:2503.10446cs.SDcs.AI2025-03被引 3

用多语言语音模型生成跨语言声纹,效果优于现有方法

Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings

  • 复用Whisper语音识别模型的编码器,结合难样本挖掘与自监督损失优化
  • 在多语种数据集上等错误率更低,AUC得分更高,尤其在非英语语种表现突出
  • 适合需要跨语言声纹识别的场景,如多语种语音分析、安全验证

多语言环境下的说话人识别面临独特挑战,尤其是传统模型主要在英语数据上训练。本文提出WSI(Whisper Speaker Identification),将预训练于大规模多语言数据的Whisper自动语音识别模型编码器重新利用,通过联合损失优化策略生成鲁棒的说话人嵌入,该策略结合在线难例三元组挖掘与自监督归一化温度交叉熵损失。借助Whisper的语言无关声学表征,该方法能有效区分不同语言和录音条件下的说话人。在多个语料库上的大量评估显示,包括VoxTube(多语言)、JVS(日语)、CallHome(德语、西班牙语、中文、日语)和Voxconverse(英语),WSI在等错误率和AUC得分上持续优于当前最优基线方法,如Pyannote Embedding、ECAPA-TDNN和Xvector。结果验证了假设:结合多语言预训练语音识别编码器与联合损失优化,可显著提升非英语语言中的说话人识别性能。

原文摘要 · Abstract (English)

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification), a framework that repurposes the encoder of the Whisper automatic speech recognition model pre trained on extensive multilingual data to generate robust speaker embeddings via a joint loss optimization strategy that leverages online hard triplet mining and self supervised Normalized Temperature-scaled Cross Entropy loss. By capitalizing on Whisper language-agnostic acoustic representations, our approach effectively distinguishes speakers across diverse languages and recording conditions. Extensive evaluations on multiple corpora, including VoxTube (multilingual), JVS (Japanese), CallHome (German, Spanish, Chinese, and Japanese), and Voxconverse (English), demonstrate that WSI consistently outperforms state-of-the-art baselines, namely Pyannote Embedding, ECAPA TDNN, and Xvector, in terms of lower equal error rates and higher AUC scores. These results validate our hypothesis that a multilingual pre-trained ASR encoder, combined with joint loss optimization, substantially improves speaker identification performance in non-English languages.

声纹识别多语言Whisper嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。