arXiv:2604.23742cs.SD2026-04ACL被引 1

构建首个实时通信语音伪造检测数据集,提升真实场景下识别能力。

RTCFake: Speech Deepfake Detection in Real-Time Communication

论文配图:RTCFake: Speech Deepfake Detection in Real-Time Communication
图 1 · 摘自论文原文
  • 提出音素引导的一致性学习策略,捕捉跨平台语义结构
  • 在未见平台和复杂噪声下仍保持高检测准确率
  • 适合语音安全、反欺诈领域研究者参考

随着语音生成技术的快速发展,实时通信(RTC)场景中的语音深度伪造威胁日益加剧。现有检测研究多集中于离线模拟,难以应对RTC传输中引入的复杂失真,包括未知的语音增强处理(如降噪)和编码器压缩。为此,我们构建了首个面向RTC场景的大规模语音伪造数据集RTCFake,总计约600小时。该数据集通过主流社交与会议平台(如Zoom)传输语音,实现离线与在线语音的精确配对。同时,提出音素引导的一致性学习(PCL)策略,使模型学习平台无关的语义结构表征。论文将RTCFake分为训练、开发和评估集,评估集包含未见的RTC平台及复杂噪声条件,提供更真实且具挑战性的评测基准。所提PCL策略显著提升跨平台泛化与噪声鲁棒性,为语音伪造检测提供了有效且通用的建模范式。数据集已公开于https://huggingface.co/datasets/JunXueTech/RTCFake。

原文摘要 · Abstract (English)

With the rapid advancement of speech generation technologies, the threat posed by speech deepfakes in real-time communication (RTC) scenarios has intensified. However, existing detection studies mainly focus on offline simulations and struggle to cope with the complex distortions introduced during RTC transmission, including unknown speech enhancement processes (e.g., noise suppression) and codec compression. To address this challenge, we present the first large-scale speech deepfake dataset tailored for RTC scenarios, termed \textit{RTCFake}, totaling approximately 600 hours. The dataset is constructed by transmitting speech through multiple mainstream social media and conferencing platforms (e.g., Zoom), enabling precise pairing between offline and online speech. In addition, we propose a phoneme-guided consistency learning (PCL) strategy that enforces models to learn platform-invariant semantic structural representations. In this paper, the RTCFake dataset is divided into training, development, and evaluation sets. The evaluation set further includes both unseen RTC platforms and unseen complex noise conditions, thereby providing a more realistic and challenging evaluation benchmark for speech deepfake detection. Furthermore, the proposed PCL strategy achieves significant improvements in both cross-platform generalization and noise robustness, offering an effective and generalizable modeling paradigm. The \textit{RTCFake} dataset is provided in the {https://huggingface.co/datasets/JunXueTech/RTCFake}.

语音伪造实时通信检测模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。