构建盲测框架评估合成语音检测模型在复杂场景下的表现。
Synthetic Audio Forensics Evaluation (SAFE) Challenge
- 设计三级挑战:原始合成、处理后音频、伪装音频,逐步提升难度。
- 覆盖90小时音频、21种真实声源、17种TTS模型,共2.1万样本。
- 为检测模型提供全面评测基准,适合研究者与安全工程师参考。
随着先进文本到语音(TTS)模型生成的合成语音越来越逼真,结合后处理和伪装技术,给音频取证检测带来了严峻挑战。本文提出SAFE(Synthetic Audio Forensics Evaluation)挑战,一个完全盲测的评估框架,用于在渐进式更难的场景下评估检测模型:原始合成语音、经过压缩/重采样等处理的音频,以及专门设计用于规避取证分析的伪装音频。SAFE挑战共包含90小时音频、21,000个音频样本,覆盖21种真实声源和17种TTS模型,分为3项任务。我们介绍了挑战设置、评估设计、数据集细节及对现有方法优劣的初步分析,为推进合成语音检测研究提供基础。更多信息请访问:https://stresearch.github.io/SAFE/
原文摘要 · Abstract (English)
The increasing realism of synthetic speech generated by advanced text-to-speech (TTS) models, coupled with post-processing and laundering techniques, presents a significant challenge for audio forensic detection. In this paper, we introduce the SAFE (Synthetic Audio Forensics Evaluation) Challenge, a fully blind evaluation framework designed to benchmark detection models across progressively harder scenarios: raw synthetic speech, processed audio (e.g., compression, resampling), and laundered audio intended to evade forensic analysis. The SAFE challenge consisted of a total of 90 hours of audio and 21,000 audio samples split across 21 different real sources and 17 different TTS models and 3 tasks. We present the challenge, evaluation design and tasks, dataset details, and initial insights into the strengths and limitations of current approaches, offering a foundation for advancing synthetic audio detection research. More information is available at \href{https://stresearch.github.io/SAFE/}{https://stresearch.github.io/SAFE/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。