用声学推理链提升伪造语音检测的可解释性与准确率
Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought
- 将低层级声学特征转为文本注入提示,引导模型逻辑推理
- 在公开数据集上比现有方法准确率更高,小模型表现优于大模型
- 适合需要可解释决策过程的语音安全应用
当前深度伪造语音检测系统多局限于二分类任务,难以生成可解释的推理过程或提供上下文丰富的解释。这类模型主要依赖隐含嵌入进行真伪判断,未能有效利用韵律、频谱和生理特征等结构化声学证据。本文提出CoLMbo-DF,一种基于特征引导的音频语言模型,通过将低层声学特征的结构化文本表示直接注入模型提示,使推理过程建立在可解释的声学证据之上,显著提升检测精度。为此,我们构建了一个包含音频对及链式思维标注的新数据集。实验表明,该方法在轻量级开源语言模型上训练,仍显著超越现有音频语言模型基线,标志着可解释深度伪造语音检测的重要进展。
原文摘要 · Abstract (English)
Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings for authenticity detection but fail to leverage structured acoustic evidence such as prosodic, spectral, and physiological attributes in a meaningful manner. This paper introduces CoLMbo-DF, a Feature-Guided Audio Language Model that addresses these limitations by integrating robust deepfake detection with explicit acoustic chain-of-thought reasoning. By injecting structured textual representations of low-level acoustic features directly into the model prompt, our approach grounds the model's reasoning in interpretable evidence and improves detection accuracy. To support this framework, we introduce a novel dataset of audio pairs paired with chain-of-thought annotations. Experiments show that our method, trained on a lightweight open-source language model, significantly outperforms existing audio language model baselines despite its smaller scale, marking a significant advancement in explainable deepfake speech detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。