arXiv:2409.10072cs.SDeess.AS2024-09中稿 · SLT被引 2

通过对比学习追踪语音转换后的原始说话人,提升验证准确性。

Speaker Contrastive Learning for Source Speaker Tracing

  • 用说话人对比损失训练嵌入提取器,增强对源说话人特征的捕捉。
  • 在挑战测试集上达到16.788%最低错误等误率,排名第一。
  • 适合对抗语音转换攻击的说话人验证系统研发者使用。

作为生物识别认证技术,说话人验证系统的安全性至关重要。然而,说话人验证系统易受各类攻击,严重影响其准确性和可靠性。其中,语音转换攻击通过改变声音特征使一人声音听起来像另一人,构成重大威胁。为应对该问题,IEEE SLT2024的源说话人追踪挑战(SSTC)旨在识别被篡改语音信号中的源说话人信息。具体而言,SSTC关注语音转换场景下的源说话人验证,判断两个转换后的语音样本是否来自同一源说话人。本文提出一种基于说话人对比学习的源说话人追踪方法,以挖掘转换语音中潜在的源说话人信息。通过在嵌入提取器训练中引入说话人对比损失,模型可从多个干扰说话人嵌入中识别出真实源说话人嵌入,从而学习更具源说话人相关性的表示。实验表明,所提方法在挑战测试集上实现16.788%的最低错误等误率(EER),获得第一名。

原文摘要 · Abstract (English)

As a form of biometric authentication technology, the security of speaker verification systems is of utmost importance. However, SV systems are inherently vulnerable to various types of attacks that can compromise their accuracy and reliability. One such attack is voice conversion, which modifies a persons speech to sound like another person by altering various vocal characteristics. This poses a significant threat to SV systems. To address this challenge, the Source Speaker Tracing Challenge in IEEE SLT2024 aims to identify the source speaker information in manipulated speech signals. Specifically, SSTC focuses on source speaker verification against voice conversion to determine whether two converted speech samples originate from the same source speaker. In this study, we propose a speaker contrastive learning-based approach for source speaker tracing to learn the latent source speaker information in converted speech. To learn a more source-speaker-related representation, we employ speaker contrastive loss during the training of the embedding extractor. This speaker contrastive loss helps identify the true source speaker embedding among several distractor speaker embeddings, enabling the embedding extractor to learn the potentially possessing source speaker information present in the converted speech. Experiments demonstrate that our proposed speaker contrastive learning system achieves the lowest EER of 16.788% on the challenge test set, securing first place in the challenge.

说话人验证语音转换对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。