arXiv:2510.24852cs.SD2025-10被引 1

用轻量卷积适配器精准识别合成语音中的多尺度时间伪影

A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection

  • 在预训练模型中加入并行多尺度卷积模块,实现高效特征提取
  • 仅用317万参数(原模型1%)即达到全微调性能
  • 适合资源受限场景下高精度语音伪造检测

近期的合成语音检测模型通常通过微调预训练的自监督学习(SSL)模型实现,但计算成本较高。参数高效微调(PEFT)提供了一种替代方案。然而,现有方法缺乏建模欺骗音频典型多尺度时间伪影所需的特定归纳偏置。本文提出多尺度卷积适配器(MultiConvAdapter),一种参数高效的架构,通过在SSL编码器中引入并行卷积模块,同时学习多个时间分辨率下的判别特征,有效捕捉短时伪影与长时失真。该方法仅需317万可训练参数(为SSL主干的1%),显著降低适应阶段的计算负担。在五个公开数据集上的评估表明,MultiConvAdapter在性能上优于全微调及主流PEFT方法。

原文摘要 · Abstract (English)

Recent synthetic speech detection models typically adapt a pre-trained SSL model via finetuning, which is computationally demanding. Parameter-Efficient Fine-Tuning (PEFT) offers an alternative. However, existing methods lack the specific inductive biases required to model the multi-scale temporal artifacts characteristic of spoofed audio. This paper introduces the Multi-Scale Convolutional Adapter (MultiConvAdapter), a parameter-efficient architecture designed to address this limitation. MultiConvAdapter integrates parallel convolutional modules within the SSL encoder, facilitating the simultaneous learning of discriminative features across multiple temporal resolutions, capturing both short-term artifacts and long-term distortions. With only $3.17$M trainable parameters ($1\%$ of the SSL backbone), MultiConvAdapter substantially reduces the computational burden of adaptation. Evaluations on five public datasets, demonstrate that MultiConvAdapter achieves superior performance compared to full fine-tuning and established PEFT methods.

语音检测参数高效卷积网络合成语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。