不听内容也能识破音频伪造,保护隐私的深度伪造检测新方法
SafeEar: Content Privacy-Preserving Audio Deepfake Detection
- 用神经音频编解码器分离语音语义与声学特征,仅用声学信息检测伪造
- 在四个数据集上等错误率降至2.02%,同时屏蔽多语言语音内容
- 适合处理涉密语音场景,兼顾隐私保护与伪造识别
文本转语音(TTS)和语音转换(VC)模型在生成逼真自然音频方面表现卓越,但其生成的音频伪造对社会和个人构成重大威胁。现有反制措施大多依赖完整原始音频判断真实性,但这些音频常含敏感内容,限制了实际应用,尤其在商业机密等敏感场景。本文提出SafeEar框架,旨在不访问音频语义内容的前提下检测深度伪造音频。核心思路是设计一种新型神经音频编解码器,将语音样本中的语义与声学信息有效解耦,仅使用声学特征(如韵律、音色)进行检测,从而避免语义内容暴露。为应对缺乏语义线索时的多样化伪造识别挑战,引入真实世界编解码增强策略。在四个基准数据集上的实验表明,SafeEar可将等错误率(EER)降至2.02%。同时,通过机器与人工听觉分析验证,五种语言的语音内容均无法被还原,词错误率(WER)均超过93.93%。此外,本文构建的用于反伪造与反内容恢复评估的基准,为音频隐私保护与伪造检测研究提供了基础。
原文摘要 · Abstract (English)
Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private content. This oversight may refrain deepfake detection from many applications, particularly in scenarios involving sensitive information like business secrets. In this paper, we propose SafeEar, a novel framework that aims to detect deepfake audios without relying on accessing the speech content within. Our key idea is to devise a neural audio codec into a novel decoupling model that well separates the semantic and acoustic information from audio samples, and only use the acoustic information (e.g., prosody and timbre) for deepfake detection. In this way, no semantic content will be exposed to the detector. To overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world codec augmentation. Extensive experiments conducted on four benchmark datasets demonstrate SafeEar's effectiveness in detecting various deepfake techniques with an equal error rate (EER) down to 2.02%. Simultaneously, it shields five-language speech content from being deciphered by both machine and human auditory analysis, demonstrated by word error rates (WERs) all above 93.93% and our user study. Furthermore, our benchmark constructed for anti-deepfake and anti-content recovery evaluation helps provide a basis for future research in the realms of audio privacy preservation and deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。