用小波散射变换提升语音深度伪造检测的可解释性与精度
WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection
- 采用小波散射变换构建多尺度、抗形变特征,融合手工特征与自监督优势
- 在Deepfake-Eval-2024上性能超越现有方法,跨数据集验证效果稳定
- 高频率和方向分辨率对捕捉细微伪造痕迹至关重要,适合可解释性研究者
本文聚焦语音深度伪造检测器的前端设计,该部分决定分类器接收的判别性声学线索。现有方法分为两类:手工滤波器组特征具有透明性但难以捕捉高层信息;自监督(SSL)特征缺乏可解释性,可能忽略细粒度频谱异常。我们提出WST-X系列新型特征提取器,通过小波散射变换(WST)级联小波卷积与模非线性,生成形变稳定、多尺度的特征。在Deepfake-Eval-2024基准测试及跨数据集评估(SpoofCeleb、In-the-Wild)中,WST-X显著优于现有前端。分析表明,较小平均尺度(J)结合高频与方向高分辨率(Q, L),对捕捉微弱伪造痕迹至关重要,凸显稳定、平移不变特征的价值。代码已开源。
原文摘要 · Abstract (English)
In this work, we focus on front-end design for speech deepfake detectors, the component that determines the discriminative acoustic cues provided to the classifier. Existing approaches are primarily categorized into two types. Hand-crafted filterbank features are transparent but limited in capturing higher-level information. SSL features, in turn, lack interpretability and may overlook fine-grained spectral anomalies. We propose the WST-X series, a novel family of feature extractors that combines the best of both worlds via the wavelet scattering transform (WST), which cascades wavelet convolutions with modulus nonlinearities to produce deformation-stable, multi-scale features. Experiments on the recent Deepfake-Eval-2024 benchmark, together with cross-dataset evaluations on the SpoofCeleb and In-the-Wild, show that WST-X outperforms existing front-ends by a wide margin. Our analysis reveals that a small averaging scale ($J$), combined with high-frequency and directional resolutions ($Q$, $L$), is critical for capturing subtle artifacts. This underscores the value of stable and translation-invariant features for speech deepfake detection. The code is available at https://github.com/xxuan-acoustics/WST-X-Series.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。