arXiv:2411.14013eess.AScs.CR2024-11被引 2

通过音频残差指纹,无需训练即可识别合成语音并溯源模型

Lightweight Model Attribution and Detection of Synthetic Speech via Audio Residual Fingerprints

  • 计算音频与滤波后版本的平均残差,提取通用指纹特征
  • 在多语言多模型下,检测与溯源准确率超99%(AUROC)
  • 适合数字取证、反欺骗等安全场景,对噪声有强鲁棒性

随着语音生成技术发展,伪造、误导和欺骗风险加剧。本文提出一种轻量级、无需训练的方法,用于检测合成语音并追溯其来源模型。该方法解决三个任务:(1)开放世界单模型溯源,(2)封闭世界多模型溯源,(3)真实与合成语音分类。核心思路是计算标准化平均残差——音频信号与其滤波版本之差——以提取不依赖模型的合成伪影指纹。在多个合成系统与语言上的实验表明,AUROC得分均高于99%,即使仅使用部分模型输出也保持可靠。方法在常见音频失真(如回声、中等背景噪声)下表现稳定,数据增强可进一步提升复杂条件下的性能。此外,通过马氏距离对比域内与域外残差指纹实现跨域检测,在未见模型上取得0.91的F1分数,验证了方法的高效性、泛化能力及在数字取证与安全中的适用性。

原文摘要 · Abstract (English)

As speech generation technologies advance, so do risks of impersonation, misinformation, and spoofing. We present a lightweight, training-free approach for detecting synthetic speech and attributing it to its source model. Our method addresses three tasks: (1) single-model attribution in an open-world setting, (2) multi-model attribution in a closed-world setting, and (3) real vs. synthetic speech classification. The core idea is simple: we compute standardized average residuals--the difference between an audio signal and its filtered version--to extract model-agnostic fingerprints that capture synthesis artifacts. Experiments across multiple synthesis systems and languages show AUROC scores above 99%, with strong reliability even when only a subset of model outputs is available. The method maintains high performance under common audio distortions, including echo and moderate background noise, while data augmentation can improve results in more challenging conditions. In addition, out-of-domain detection is performed using Mahalanobis distances to in-domain residual fingerprints, achieving an F1 score of 0.91 on unseen models, reinforcing the method's efficiency, generalizability, and suitability for digital forensics and security applications.

语音检测合成语音数字取证残差指纹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。