arXiv:2506.02590cs.SDcs.CL2025-06被引 10

用度量学习追踪合成语音来源,提升伪造音频溯源能力。

Synthetic Speech Source Tracing using Metric Learning

  • 采用度量学习与分类对比方法,从语音中识别生成模型
  • 在MLAADv5上,ResNet性能媲美甚至超过自监督模型
  • 适合音频取证与反伪造技术研究者参考

本文针对合成语音的来源追溯问题,提出基于说话人识别思路的解决方案。现有工作多聚焦于欺骗检测,而来源追溯缺乏稳健方法。我们评估了两类策略:基于分类与度量学习的方法,并在MLAADv5基准上使用ResNet和自监督学习(SSL)主干网络进行测试。结果表明,ResNet在度量学习框架下表现优异,性能可媲美甚至超越部分基于SSL的系统。研究验证了ResNet在该任务中的可行性,同时指出需优化SSL表征以更好适配来源追溯。本工作将说话人识别方法引入音频取证,为对抗合成媒体操纵提供新方向。

原文摘要 · Abstract (English)

This paper addresses source tracing in synthetic speech-identifying generative systems behind manipulated audio via speaker recognition-inspired pipelines. While prior work focuses on spoofing detection, source tracing lacks robust solutions. We evaluate two approaches: classification-based and metric-learning. We tested our methods on the MLAADv5 benchmark using ResNet and self-supervised learning (SSL) backbones. The results show that ResNet achieves competitive performance with the metric learning approach, matching and even exceeding SSL-based systems. Our work demonstrates ResNet's viability for source tracing while underscoring the need to optimize SSL representations for this task. Our work bridges speaker recognition methodologies with audio forensic challenges, offering new directions for combating synthetic media manipulation.

语音溯源度量学习音频取证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。