实时检测RVC语音转换,防伪克隆语音欺骗
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
- 用短窗音频特征捕捉RVC语音的细微差异
- 在嘈杂背景中仍能准确识别伪造语音
- 适合通信安全、反诈骗场景使用
生成式音频技术现已实现高度逼真的语音克隆与实时语音转换,增加了电话和视频通话等通信渠道中的冒名顶替、欺诈和虚假信息风险。本研究针对基于检索的语音转换(RVC)生成的语音,评估其在DEEP-VOICE数据集上的实时检测能力,该数据集包含多位知名人士的真实与转换语音样本。为模拟真实环境,对孤立声学成分进行深度伪造处理,并重新引入背景音以抑制明显伪影,突出转换特有的线索。将检测任务建模为流式分类,将音频划分为1秒片段,提取时频与倒谱特征,训练监督学习模型判断每段为真实或转换语音。所提系统支持低延迟推理,可实现段级判定与通话级聚合。实验表明,短窗声学特征能可靠捕获与RVC语音相关的判别模式,即使在噪声环境中亦然。结果验证了实用化实时深度伪造语音检测的可行性,并强调在真实音频混音条件下评估的重要性以确保鲁棒部署。
原文摘要 · Abstract (English)
Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study investigates real-time detection of AI-generated speech produced using Retrieval-based Voice Conversion (RVC), evaluated on the DEEP-VOICE dataset, which includes authentic and voice-converted speech samples from multiple well-known speakers. To simulate realistic conditions, deepfake generation is applied to isolated vocal components, followed by the reintroduction of background ambiance to suppress trivial artifacts and emphasize conversion-specific cues. We frame detection as a streaming classification task by dividing audio into one-second segments, extracting time-frequency and cepstral features, and training supervised machine learning models to classify each segment as real or voice-converted. The proposed system enables low-latency inference, supporting both segment-level decisions and call-level aggregation. Experimental results show that short-window acoustic features can reliably capture discriminative patterns associated with RVC speech, even in noisy backgrounds. These findings demonstrate the feasibility of practical, real-time deepfake speech detection and underscore the importance of evaluating under realistic audio mixing conditions for robust deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。