让语音伪造检测结果可追溯,提升可信度与审查效率。
From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection
- 用四类证据构建可审计的决策记录,保留判断依据。
- 在ASVspoof数据集上将误报率降至8.43%,低于仅依赖检索的方法。
- 适合需要透明审查的高风险语音验证场景使用。
语音深度伪造能以高度逼真的方式模仿说话人声音,欺骗人类听者和自动化系统。尽管检测技术进步显著,但现有检测器通常只输出单个评分,难以解释为何某些临界样本应被信任、延迟或审查。两个样本可能处于同一评分区间,却因被动检测与检索证据冲突,或关键探测器不可用等原因所致。本文提出一种可审计决策记录,在不放弃单一评分的前提下,整合被动检测得分、带标记衍生品的条件性密钥探测得分、检索支持度及说话人档案差距,并显式标注分歧坐标。在包含4,080个样本的ASVspoof 5 Track 1匹配子集上,结合检索增强的固定规则使等错误率(EER)从15.84%降至11.91%;后期校准后进一步降至8.43%。当审查预算为33.75%时,暴露的线索联合覆盖了校准模型82.85%的错误。最优被动WavLM模型仍达6.71% EER,因此本方法不作为更强的独立检测器,而是保留每个被识别样本背后的证据,同时维持单一可阈值化评分用于审查决策。
原文摘要 · Abstract (English)
Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress in speech deepfake detection, but most detectors still end with one score per utterance. That score is useful for ranking systems, yet it says little about why a borderline item should be trusted, deferred, or reviewed. Two utterances can fall in the same score band for different reasons, for example because passive and retrieval evidence disagree or because the keyed probe is unavailable. We ask whether the final decision can remain scalar without discarding that provenance. We answer this question with an auditable decision record that carries four aligned cues into a late calibration step: a passive detector score, a conditional keyed-probe score on a marked derivative, retrieval support, and a speaker-profile margin, together with explicit disagreement coordinates. On the 4,080-example ASVspoof 5 Track 1 matched subset, the fixed retrieval-augmented rule improves on retrieval-only evidence, from 15.84 percent to 11.91 percent EER, and late calibration over the full record reaches 8.43 percent EER. At a 33.75 percent review budget, the exposed cue union covers 82.85 percent of the calibrated model's errors. The best passive WavLM run still reaches 6.71 percent EER, so we do not present the decision record as a stronger standalone detector. Its contribution is to preserve the evidence behind each surfaced utterance while still producing one operating score for thresholding and review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。