arXiv:2601.23066cs.SDcs.AI2026-01被引 7

让语音大模型更关注声学细节,提升深伪语音检测精度

Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection

  • 融合原始音频与结构化频谱图,显式增强声学特征感知
  • 在语义误导场景下检测准确率显著提升,尤其对细微声学异常敏感
  • 适合需要高精度语音安全检测的研究与应用者

语音深伪检测(SDD)旨在判断给定语音信号是真实还是合成生成。现有基于音频大语言模型(LLM)的方法擅长内容理解,但其预测常受语义相关线索干扰,导致细粒度声学痕迹被忽略。因此,即使存在微弱声学异常,语义自然的伪造语音仍可能绕过检测。这表明问题并非缺乏声学数据,而是语义主导推理时声学信息难以获取。为此,本文在音频LLM范式下研究SDD,提出增强听觉感知的音频大语言模型框架(SDD-APALLM),通过结合原始音频与结构化频谱图,显式暴露时间-频率层面的细粒度声学证据,使模型能有效捕捉细微声学不一致,同时保持语义理解能力。实验表明,该方法在检测准确率和鲁棒性上均有持续提升,尤其在语义误导场景下表现优异。进一步分析显示,性能提升源于语义与声学信息的协同利用,而非简单的模态叠加。

原文摘要 · Abstract (English)

Speech deepfake detection (SDD) focuses on identifying whether a given speech signal is genuine or has been synthetically generated. Existing audio large language model (LLM)-based methods excel in content understanding; however, their predictions are often biased toward semantically correlated cues, which results in fine-grained acoustic artifacts being overlooked during the decisionmaking process. Consequently, fake speech with natural semantics can bypass detectors despite harboring subtle acoustic anomalies; this suggests that the challenge stems not from the absence of acoustic data, but from its inadequate accessibility when semantic-dominant reasoning prevails. To address this issue, we investigate SDD within the audio LLM paradigm and introduce SDD with Auditory Perception-enhanced Audio Large Language Model (SDD-APALLM), an acoustically enhanced framework designed to explicitly expose fine-grained time-frequency evidence as accessible acoustic cues. By combining raw audio with structured spectrograms, the proposed framework empowers audio LLMs to more effectively capture subtle acoustic inconsistencies without compromising their semantic understanding. Experimental results indicate consistent gains in detection accuracy and robustness, especially in cases where semantic cues are misleading. Further analysis reveals that these improvements stem from a coordinated utilization of semantic and acoustic information, as opposed to simple modality aggregation.

语音检测深度伪造音频大模型声学感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。