通过呼吸线索提升音频深度伪造检测的泛化能力
BreathNet: Generalizable Audio Deepfake Detection via Breath-Cue-Guided Feature Refinement
- 利用呼吸声特征引导时序特征增强,捕捉细微生理线索
- 在多个数据集上实现最低错误率,最差场景下EER仅4.94%
- 适合需要高鲁棒性的语音安全系统开发者使用
随着深度伪造音频愈发逼真多样,开发具备泛化能力的防御系统变得至关重要。现有方法主要依赖XLS-R前端特征提升泛化性,但仍受限于对细粒度信息(如生理线索或频域特征)关注不足。本文提出BreathNet,一种融合细粒度呼吸信息的音频深度伪造检测框架。设计BreathFiLM机制,根据呼吸声存在情况选择性放大时序表示,并与XLS-R提取器联合训练,促使提取器学习并编码呼吸相关线索。同时,采用频域前端提取谱特征,与时序特征融合以补充声码器或压缩伪影带来的互补信息。此外,提出包含正样本监督对比损失(PSCL)、中心损失和对比损失的一组特征损失,共同增强模型判别能力,使真实与伪造样本在特征空间中更易区分。在五个基准数据集上的大量实验表明,该方法达到当前最优性能:使用ASVspoof 2019 LA训练集,在四个相关评测集上平均等错误率(EER)为1.99%,尤其在In-the-Wild数据集上达4.70% EER;在ASVspoof5评估协议下,最新基准上取得4.94% EER。
原文摘要 · Abstract (English)
As deepfake audio becomes more realistic and diverse, developing generalizable countermeasure systems has become crucial. Existing detection methods primarily depend on XLS-R front-end features to improve generalization. Nonetheless, their performance remains limited, partly due to insufficient attention to fine-grained information, such as physiological cues or frequency-domain features. In this paper, we propose BreathNet, a novel audio deepfake detection framework that integrates fine-grained breath information to improve generalization. Specifically, we design BreathFiLM, a feature-wise linear modulation mechanism that selectively amplifies temporal representations based on the presence of breathing sounds. BreathFiLM is trained jointly with the XLS-R extractor, in turn encouraging the extractor to learn and encode breath-related cues into the temporal features. Then, we use the frequency front-end to extract spectral features, which are then fused with temporal features to provide complementary information introduced by vocoders or compression artifacts. Additionally, we propose a group of feature losses comprising Positive-only Supervised Contrastive Loss (PSCL), center loss, and contrast loss. These losses jointly enhance the discriminative ability, encouraging the model to separate bona fide and deepfake samples more effectively in the feature space. Extensive experiments on five benchmark datasets demonstrate state-of-the-art (SOTA) performance. Using the ASVspoof 2019 LA training set, our method attains 1.99% average EER across four related eval benchmarks, with particularly strong performance on the In-the-Wild dataset, where it achieves 4.70% EER. Moreover, under the ASVspoof5 evaluation protocol, our method achieves an EER of 4.94% on this latest benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。