arXiv:2410.09869cs.SDcs.AI2024-10中稿 · Interspeech 2024被引 10

用提示调优提升音频伪造检测在小数据下的适应性与效率

Prompt Tuning for Audio Deepfake Detection: Computationally Efficient Test-time Domain Adaptation with Limited Target Dataset

  • 通过提示调优无缝融合主流模型,缓解源域与目标域差异
  • 仅需少量目标数据即可有效适配,且不增加额外参数量
  • 计算开销极低,适合资源受限场景下的实时检测

我们研究音频深度伪造检测(ADD)中的测试时领域自适应问题,解决三大挑战:(i) 源域与目标域之间的分布差异,(ii) 目标数据集规模有限,(iii) 高计算成本。提出一种基于提示调优的插件式ADD方法,可无缝集成于当前主流Transformer模型或其他微调方法中,有效缓解域间差异并提升目标数据上的性能(挑战(i))。该方法因无需大量新增参数,可高效适应小规模目标数据集(挑战(ii)),同时显著降低计算开销,克服了大模型在ADD任务中常见的高资源消耗问题(挑战(iii))。结果表明,在存在域差异的情况下,提示调优为实现高精度检测提供了低数据、低开销的可行路径。

原文摘要 · Abstract (English)

We study test-time domain adaptation for audio deepfake detection (ADD), addressing three challenges: (i) source-target domain gaps, (ii) limited target dataset size, and (iii) high computational costs. We propose an ADD method using prompt tuning in a plug-in style. It bridges domain gaps by integrating it seamlessly with state-of-the-art transformer models and/or with other fine-tuning methods, boosting their performance on target data (challenge (i)). In addition, our method can fit small target datasets because it does not require a large number of extra parameters (challenge (ii)). This feature also contributes to computational efficiency, countering the high computational costs typically associated with large-scale pre-trained models in ADD (challenge (iii)). We conclude that prompt tuning for ADD under domain gaps presents a promising avenue for enhancing accuracy with minimal target data and negligible extra computational burden.

音频伪造检测提示调优领域自适应高效检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。