arXiv:2607.28351cs.SDcs.AI2026-07

Teffic-Audio 用简单架构+精心训练,高效识别多种语音伪造。

Teffic-Audio: Tell Fact from Fiction

论文配图:Teffic-Audio: Tell Fact from Fiction
图 1 · 摘自论文原文
  • 基于Conformer的简单结构,通过数据多样性与采样平衡提升泛化能力。
  • 在14个测试集上均值等错误率仅1.454%,优于所有公开系统。
  • 适合需要高鲁棒性且资源受限的语音伪造检测场景。

语音深度伪造检测面临越来越多样化的伪造手段,包括语音合成、声纹转换、声码器重建和神经编码重合成。这些伪造痕迹还受源语音、录音环境和传输通道差异的影响。因此,跨异构条件的鲁棒泛化成为实际检测系统的核心要求。本文提出Teffic-Audio,一个面向全面评估的通用语音深度伪造检测系统。该系统采用基于Conformer的语音编码器、多头注意力统计池化和二分类器的简洁架构。不依赖复杂结构,而是通过多源数据融合、攻击与源平衡采样及多样化音频增强来提升泛化性能。仅使用开源数据训练,Teffic-Audio在Speech-DF-Arena的14个测试集上达到1.454%的均值等错误率(EER),超越所有现有公开系统。其在5个独立测试集上也取得最低EER,且相比更大规模领先系统具有更优的性能-复杂度权衡。整体上,Teffic-Audio为通用语音深度伪造检测提供了强有力的实用基准。

原文摘要 · Abstract (English)

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-Audio, a general speech deepfake detection system designed for comprehensive evaluation environment. Teffic-Audio adopts a straightforward detector architecture consisting of a Conformer-based speech encoder, multi-head attentive statistics pooling, and a binary classifier. Rather than relying on additional architectural complexity, the system improves generalization through its training recipe, which integrates multi-source data, attack- and source-balanced sampling, and diverse audio augmentation. Trained only with open-source data, Teffic-Audio achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard. It also obtains the lowest EER on five individual test sets and shows a favorable performance-complexity trade-off compared with larger leading systems. Overall, Teffic-Audio provides a strong and practical reference system for general speech deepfake detection.

语音伪造深度伪造检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。