arXiv:2603.06164cs.SDcs.AI2026-03被引 7

小模型也能抗干扰:多语言预训练让紧凑模型媲美大模型

Do Compact SSL Backbones Matter for Audio Deepfake Detection? A Controlled Study with RAPTOR

  • 用统一融合框架对比HuBERT与WavLM等紧凑模型
  • 100M参数模型通过多语言预训练达到商用系统水平
  • 发现模型鲁棒性来自预训练路径而非模型大小

自监督学习(SSL)是现代语音深度伪造检测的基础,但以往研究多集中于单一大型wav2vec2-XLSR骨干网络,对紧凑模型关注不足。本文提出RAPTOR——一种面向跨域识别的表征感知成对门控变压器,并在统一的成对门控融合检测器中,对HuBERT和WavLM系列紧凑模型进行了控制性研究,覆盖14个跨域基准。结果表明,多语言HuBERT预训练是跨域鲁棒性的主要驱动力,使100M参数模型性能媲美更大规模及商用系统。除误报率(EER)外,本文引入基于扰动的测试时增强协议与随机不确定性分析,揭示标准指标无法捕捉的校准差异:WavLM变体在扰动下表现出过度自信的校准偏差,而迭代式mHuBERT则保持稳定。研究结论指出,SSL预训练轨迹而非模型规模,才是可靠语音深度伪造检测的关键。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) underpins modern audio deepfake detection, yet most prior work centers on a single large wav2vec2-XLSR backbone, leaving compact under studied. We present RAPTOR, Representation Aware Pairwise-gated Transformer for Out-of-domain Recognition a controlled study of compact SSL backbones from the HuBERT and WavLM within a unified pairwise-gated fusion detector, evaluated across 14 cross-domain benchmarks. We show that multilingual HuBERT pre-training is the primary driver of cross-domain robustness, enabling 100M models to match larger and commercial systems. Beyond EER, we introduce a test-time augmentation protocol with perturbation-based aleatoric uncertainty to expose calibration differences invisible to standard metrics: WavLM variants exhibit overconfident miscalibration under perturbation, whereas iterative mHuBERT remains stable. These findings indicate that SSL pre-training trajectory, not model scale, drives reliable audio deepfake detection.

音频伪造检测自监督学习模型压缩鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。