arXiv:2601.03615cs.CLcs.SD2026-01被引 2

检测音频伪造时,模型的推理过程是否可靠?

SARA: Stress Test Reasoning in Audio Deepfake Detection

  • 提出SARA框架,从听觉感知、推理与判断一致性等维度评估语音模型推理可靠性
  • 声学攻击使推理与判断一致性平均下降14.20%,引发内部逻辑矛盾
  • 推理文本的连贯性可作隐蔽信号,无需原始音频即可检测干扰(F1=0.78)

语音语言模型(ALMs)为可解释的音频深度伪造检测(ADD)提供了新方向,通过推理轨迹实现预测透明化。然而,这些推理可能缺乏实际支持,表现为不连贯或用看似合理却误导性的解释来合理化错误判断。此外,现有研究对ALM在对抗攻击下的行为尚缺乏深入探索,其解释能力的实际可靠性存疑。为此,本文提出SARA(Shift Analysis of Reasoning in Audio),一个从声学感知、推理-判断一致性及矛盾性三个维度评估ALM推理表现的诊断框架。我们测试了五种开源ALMs在声学与语言对抗攻击下的表现。结果表明,声学攻击显著降低推理-判断一致性(平均下降14.20%),常引发内部逻辑冲突;而语言攻击虽成功率更高,但能较好维持推理连贯性。进一步发现,生成推理轨迹的文本连贯性可作为对抗输入的潜在指示器,在不访问原始音频信号的情况下,仍能有效检测干扰样本(F1=0.78)。这说明,即便最终分类结果被破坏,推理轨迹仍具有诊断价值。

原文摘要 · Abstract (English)

Audio Language Models (ALMs) offer a promising shift towards explainable audio deepfake detections (ADD), moving beyond \textit{black-box} classifiers by providing transparency to their predictions via reasoning traces. However, such reasoning may not support the model predictions, reflecting poor coherence, or, worse, may rationalize incorrect predictions with plausible but misleading explanation. Moreover, the behavior of ALM reasoning under adversarial attacks remains under-explored, raising questions about the practical reliability of such explanation capabilities. To address this gap, this study introduces \textbf{SARA} (\textbf{S}hift \textbf{A}nalysis of \textbf{R}easoning in \textbf{A}udio), a diagnostic framework that evaluates ALM reasoning across three dimensions: acoustic perception, reasoning-verdict coherence and dissonance. We test five open-source ALMs against both acoustic and linguistic adversarial attacks. We show that acoustic attacks significantly degrade reasoning-verdict coherence (average decrease of 14.20\%), frequently inducing internal logical conflicts. Conversely, linguistic attacks achieve higher attack success rates while maintaining reasoning coherence. We further demonstrate that the textual coherence of generated reasoning traces also serves as a latent indicator of adversarial inputs, enabling effective detection of perturbed audio (0.78 in F1) \textit{without accessing the raw acoustic signal}. These findings suggest that reasoning traces provide diagnostic utility that persists even when final classification outputs are compromised.

音频伪造可解释性对抗攻击推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。