提出时频一致性学习框架,提升语音深度伪造检测在真实音频处理链下的鲁棒性。
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

- 设计时频一致约束,学习经声学前端处理前后仍稳定的欺骗特征。
- 在真实前端处理链下,检测准确率提升12.3%,显著缓解性能下降。
- 适合关注实际部署中语音伪造检测鲁棒性的研究者与工程师。
近年来,语音深度伪造检测(SDD)取得显著进展,但其鲁棒性评估仍局限于受控加性噪声场景,缺乏对真实部署中声学前端(AFE)处理流程引入复杂失真的系统研究。本文模拟包含回声消除、降噪、自动增益控制和语音活动检测(VAD)的统一AFE流水线,对当前主流模型进行全面评估。结果表明,AFE引入的非线性与时频耦合失真会显著降低检测性能。为此,提出时频一致性学习(TFCL)框架,旨在学习在AFE处理前后保持稳定的欺骗表征。发现AFE不仅造成时间错位(如VAD导致的片段级偏移),还弱化或扭曲关键频域特征。TFCL采用注意力驱动的软对齐机制捕捉跨时序依赖,并引入频域结构一致性约束以强化特征不变性。实验表明,该方法能有效抑制AFE带来的性能退化,在多种真实场景下显著提升SDD鲁棒性。代码已开源。
原文摘要 · Abstract (English)
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo cancellation, noise suppression, automatic gain control, and voice activity detection (VAD), and conduct a comprehensive evaluation of current state-of-the-art models. The results show that the nonlinear and time-frequency coupled distortions introduced by AFE significantly degrade detection performance. To address this issue, we propose a Time-Frequency Consistency Learning (TFCL) framework, which aims to learn invariant spoofing representations that remain stable before and after AFE processing. We observe that AFE not only introduces temporal misalignment (e.g., segment-level shifts caused by VAD), but also weakens or distorts critical frequency-domain cues. To this end, TFCL employs an attention-driven soft alignment mechanism to capture cross-temporal dependencies, along with frequency-domain structural consistency constraints to enforce feature invariance. As a result, the model is able to maintain stable representations under both temporal perturbations and spectral distortions. Extensive experimental results demonstrate that the proposed method effectively mitigates the performance degradation caused by AFE processing, significantly improving the robustness of SDD in real-world scenarios. The code is available at https://github.com/JunXue-tech/TFCL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。