arXiv:2603.21875eess.AScs.CL2026-03中稿 · Interspeech 2026被引 1

提出新方法分离说话人特征,提升深度伪造语音溯源准确性。

Disentangling Speaker Traits for Deepfake Source Verification via Chebyshev Polynomial and Riemannian Metric Learning

  • 用切比雪夫多项式缓解解耦优化时的梯度不稳问题。
  • 将语音与说话人嵌入投影到双曲空间,降低说话人信息干扰。
  • 在新评测协议下显著提升伪造语音源识别效果,适合安全验证场景。

语音深度伪造溯源系统旨在判断两段合成语音是否来自同一生成器,通常假设源嵌入与说话人特征无关,但该假设尚未验证。本文首次研究说话人因素对溯源的影响,提出一种说话人解耦度量学习(SDML)框架,包含两个新损失函数:其一利用切比雪夫多项式缓解解耦优化中的梯度不稳定性;其二将源与说话人嵌入投影至双曲空间,通过黎曼度量距离减少说话人信息,学习更具判别性的源特征。在MLAAD基准上,基于四个新提出的解耦评测协议进行评估,结果验证了SDML框架的有效性。代码、评测协议及演示网站已公开于https://github.com/xxuan-acoustics/RiemannSD-Net。

原文摘要 · Abstract (English)

Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting source embeddings are independent of speaker traits. However, this assumption remains unverified. In this paper, we first investigate the impact of speaker factors on source verification. We propose a speaker-disentangled metric learning (SDML) framework incorporating two novel loss functions. The first leverages Chebyshev polynomial to mitigate gradient instability during disentanglement optimization. The second projects source and speaker embeddings into hyperbolic space, leveraging Riemannian metric distances to reduce speaker information and learn more discriminative source features. Experimental results on MLAAD benchmark, evaluated under four newly proposed protocols designed for source-speaker disentanglement scenarios, demonstrate the effectiveness of SDML framework. The code, evaluation protocols and demo website are available at https://github.com/xxuan-acoustics/RiemannSD-Net.

语音伪造解耦学习双曲空间度量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。