提出新方法让假音频检测模型不再依赖说话人特征,提升跨数据集识别能力。
Beyond Identity: A Generalizable Approach for Deepfake Audio Detection
- 用分离合成痕迹的模块,让模型专注识别伪造特征而非说话人身份。
- 在多个数据集上测试,最高F1达0.813,显著优于基线模型。
- 适合需要高泛化能力的音频安全检测场景,如反诈骗系统。
假音频对数字安全构成日益严重的威胁,可能用于社会工程、欺诈和身份盗用。然而,现有检测模型在跨数据集时泛化能力差,原因在于隐式身份泄漏——模型无意中学习了说话人特有特征而非伪造痕迹。据我们所知,这是首个在音频假音频检测领域显式分析并解决身份泄漏的研究。本文提出一种身份无关的假音频检测框架,通过鼓励模型关注伪造特定痕迹而非过度拟合说话人特征,缓解身份泄漏问题。方法采用人工痕迹检测模块(ADMs),在时域和频域分离合成痕迹,增强跨数据集泛化能力。引入新型动态痕迹生成技术,包括频域交换、时域操作和背景噪声增强,以强化对数据集不变特征的学习。在ASVspoof2019、ADD 2022、FoR和In-The-Wild数据集上的大量实验表明,基于ADM的模型在ADD 2022、FoR和In-The-Wild上的F1得分分别达到0.230、0.604和0.813,持续优于基线。动态频域交换在多种条件下表现最佳。研究强调了基于痕迹学习在缓解隐式身份泄漏、提升泛化能力方面的价值。
原文摘要 · Abstract (English)
Deepfake audio presents a growing threat to digital security, due to its potential for social engineering, fraud, and identity misuse. However, existing detection models suffer from poor generalization across datasets, due to implicit identity leakage, where models inadvertently learn speaker-specific features instead of manipulation artifacts. To the best of our knowledge, this is the first study to explicitly analyze and address identity leakage in the audio deepfake detection domain. This work proposes an identity-independent audio deepfake detection framework that mitigates identity leakage by encouraging the model to focus on forgery-specific artifacts instead of overfitting to speaker traits. Our approach leverages Artifact Detection Modules (ADMs) to isolate synthetic artifacts in both time and frequency domains, enhancing cross-dataset generalization. We introduce novel dynamic artifact generation techniques, including frequency domain swaps, time domain manipulations, and background noise augmentation, to enforce learning of dataset-invariant features. Extensive experiments conducted on ASVspoof2019, ADD 2022, FoR, and In-The-Wild datasets demonstrate that the proposed ADM-enhanced models achieve F1 scores of 0.230 (ADD 2022), 0.604 (FoR), and 0.813 (In-The-Wild), consistently outperforming the baseline. Dynamic Frequency Swap proves to be the most effective strategy across diverse conditions. These findings emphasize the value of artifact-based learning in mitigating implicit identity leakage for more generalizable audio deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。