通过概率属性嵌入解析伪造语音合成组件,实现可解释的检测与溯源。
Towards Explainable Spoofed Speech Attribution and Detection:a Probabilistic Approach for Characterizing Speech Synthesizer Components
- 将伪造语音分解为可解释的概率属性嵌入,识别合成关键组件。
- 检测任务平衡准确率99.7%,溯源任务达90.23%,接近原始嵌入性能。
- 结合Shapley值量化各属性贡献,适合需透明决策的安防场景。
我们提出一种可解释的概率框架,通过将伪造语音分解为概率属性嵌入来表征其特征。与缺乏可解释性的高维反欺诈嵌入不同,该框架利用高层属性及其取值,精准识别语音合成器的具体组件。采用四种分类器后端,解决两个下游任务:伪造语音检测(真/假判别)与攻击溯源(识别生成方法)。引入Shapley值量化各属性对决策的相对贡献。在ASVspoof2019数据集上,持续时长和转换建模对检测至关重要;波形生成与说话人建模则主导溯源。检测任务中,该嵌入达99.7%平衡准确率与0.22%等错误率,接近原始嵌入(99.9%与0.22%);溯源任务中,分别达到90.23%与2.07%,与原始嵌入(90.16%与2.11%)相当。结果表明,该框架兼具内在可解释性与高性能。
原文摘要 · Abstract (English)
We propose an explainable probabilistic framework for characterizing spoofed speech by decomposing it into probabilistic attribute embeddings. Unlike raw high-dimensional countermeasure embeddings, which lack interpretability, the proposed probabilistic attribute embeddings aim to detect specific speech synthesizer components, represented through high-level attributes and their corresponding values. We use these probabilistic embeddings with four classifier back-ends to address two downstream tasks: spoofing detection and spoofing attack attribution. The former is the well-known bonafide-spoof detection task, whereas the latter seeks to identify the source method (generator) of a spoofed utterance. We additionally use Shapley values, a widely used technique in machine learning, to quantify the relative contribution of each attribute value to the decision-making process in each task. Results on the ASVspoof2019 dataset demonstrate the substantial role of duration and conversion modeling in spoofing detection; and waveform generation and speaker modeling in spoofing attack attribution. In the detection task, the probabilistic attribute embeddings achieve $99.7\%$ balanced accuracy and $0.22\%$ equal error rate (EER), closely matching the performance of raw embeddings ($99.9\%$ balanced accuracy and $0.22\%$ EER). Similarly, in the attribution task, our embeddings achieve $90.23\%$ balanced accuracy and $2.07\%$ EER, compared to $90.16\%$ and $2.11\%$ with raw embeddings. These results demonstrate that the proposed framework is both inherently explainable by design and capable of achieving performance comparable to raw CM embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。