检测伪造音频时,不依赖内容与说话人特征,准确率稳定在99%以上。
DETECT-3B-Omni is Agnostic of Content and Demographics
- 通过等效性检验验证检测器对语义和人口特征的无关性。
- 在30个州、8种克隆系统下,准确率差异不超过2个百分点。
- 适合需要公平合规的深度伪造音频检测场景。
可信且符合GDPR要求的深度伪造音频检测器必须基于声学伪影做决策,而非所讲内容或说话人身份。我们对Resemble AI的DETECT-3B-Omni检测器进行了大规模语义独立性研究。使用来自美国30个州、涵盖不同性别与年龄的10,240个音频样本,这些样本由8种不同的AI语音克隆系统生成,测试检测准确率是否受语义内容(良性/恶意)、说话人性别、年龄或地区影响。通过等效性检验,结果表明任意两组间的准确率差异在99%置信度下不超过2个百分点。因此,该检测器在识别AI生成音频时,对内容与说话人特征具有同等准确性。
原文摘要 · Abstract (English)
A trustworthy and GDPR-compliant deepfake audio detector must base its decisions on acoustic artifacts, not on what is being said or who is speaking. We present a large-scale study of semantic independence for Resemble AI's detector, DETECT-3B-Omni. Using 10,240 audio samples from diverse US English speakers across 30 states, generated through 8 different AI voice-cloning systems, we test whether detection accuracy depends on spoken content (benign versus malicious), speaker gender, speaker age, or speaker region. Using equivalence testing, our results show that the accuracy difference between any two of these groups is at most 2 percentage points, at 99% confidence. The detector therefore identifies AI-generated audio with equivalent accuracy regardless of what the audio says or who the speaker is.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。