提出新方法诊断语音伪造检测器对说话人身份的依赖问题。
Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

- 用身份敏感度评分(ISS)量化检测器对说话人变化的响应程度。
- 错误分类样本的ISS值比正确样本高29至52倍,预测准确率AUC达0.954。
- 适合研究语音伪造检测漏洞或提升模型鲁棒性的研究人员使用。
语音深伪检测器在标准数据集上表现良好,但在不同数据集上错误率可能上升二十倍。我们指出,一个原因是检测器过度依赖说话人身份:训练数据中说话人身份与真实/合成标签相关,导致检测器部分依赖说话人线索而非合成痕迹。为此提出身份敏感度评分(ISS),一种无需推理时真值标签的逐句诊断工具,通过检测器输出与参考说话人样本计算。在两个检测器和两个数据集上,误判样本的ISS值比正确样本高29至52倍,且ISS单独可实现最高0.954的AUC预测。为验证ISS是否真正反映身份敏感性,对500个语音样本进行语音转换,结果发现被ISS标记为敏感的样本响应强度是稳定样本的19至30倍。这表明ISS能有效识别检测器的说话人依赖性缺陷,适用于推理阶段的身份敏感性分析。
原文摘要 · Abstract (English)
Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different dataset. We argue that one contributing factor is speaker-identity reliance: standard training corpora correlate speaker identity with the genuine/synthetic label, allowing detectors to partially rely on speaker-related cues rather than synthesis artifacts alone. We propose the Identity Sensitivity Score (ISS), a per-utterance diagnostic that quantifies how much a detector's output changes across different speaker identity contexts. ISS requires no ground-truth labels at inference time and can be computed from the detector score and a pool of reference speaker examples. Across two detectors and two datasets, incorrectly classified utterances have ISS scores 29 to 52 times higher than correctly classified utterances, and ISS alone predicts misclassification with area-under-curve (AUC) up to 0.954. To test whether ISS actually captures identity-sensitive behavior rather than serving only as a proxy for prediction confidence, we apply voice conversion to 500 utterances and measure the resulting detector-score shift. Utterances flagged as identity-sensitive by ISS respond 19 to 30 times more strongly to this manipulation than utterances flagged as stable. These results position ISS as a practical inference-time diagnostic for speaker-dependent failure analysis in audio deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。