arXiv:2512.08403cs.SD2025-12被引 9

通过优化音频大模型组件,实现跨任务通用的深度伪造检测

DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components

  • 设计可泛化的音频大模型结构,融合音视频编码与文本LLM
  • 在多个数据集上达到95.76%平均准确率,超越现有方法
  • 适合需要多任务检测与跨域适应的研究者和安全应用

音频深度伪造检测因对安全与可信性的潜在威胁而备受关注。传统深度学习方法虽广泛应用,但在面对新出现的伪造技术及更复杂的任务(如伪造归属识别)时泛化能力不足。尽管大语言模型(LLMs)理论上具备强泛化能力,但现有音频大模型(ALLMs)在深度伪造检测中仍存在性能瓶颈,即使数据充足亦然。本研究分析了ALLM的核心组件——音频编码器与文本型大模型的影响,发现其选择与组合方式至关重要。我们提出一种新型ALLM架构,能有效泛化至域外伪造测试及其他深度伪造任务(如伪造定位与归属识别)。实验表明,该模型在ASVSpoof2019、InTheWild和Demopage等多个数据集上达到领先性能,平均准确率达95.76%,且在归属与定位等任务中表现优于当前主流音频理解模型。数据与代码已附于补充材料。

原文摘要 · Abstract (English)

Audio deepfake detection has recently garnered public concern due to its implications for security and reliability. Traditional deep learning methods have been widely applied to this task but often lack generalisability when confronted with newly emerging spoofing techniques and more tasks such as spoof attribution recognition rather than simple binary classification. In principle, Large Language Models (LLMs) are considered to possess the needed generalisation capabilities. However, previous research on Audio LLMs (ALLMs) indicates a generalization bottleneck in audio deepfake detection performance, even when sufficient data is available. Consequently, this study investigates the model architecture and examines the effects of the primary components of ALLMs, namely the audio encoder and the text-based LLM. Our experiments demonstrate that the careful selection and combination of audio encoders and text-based LLMs are crucial for unlocking the deepfake detection potential of ALLMs. We further propose an ALLM structure capable of generalizing deepfake detection abilities to out-of-domain spoofing tests and other deepfake tasks, such as spoof positioning and spoof attribution recognition. Our proposed model architecture achieves state-of-the-art (SOTA) performance across multiple datasets, including ASVSpoof2019, InTheWild, and Demopage, with accuracy reaching up to 95.76% on average, and exhibits competitive capabilities in other deepfake detection tasks such as attribution, and localisation compared to SOTA audio understanding models. Data and codes are provided in supplementary materials.

深度伪造检测音频大模型多任务泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。