arXiv:2606.16137cs.CLcs.AI2026-06中稿 · Interspeech 2026

用XAI证据增强大模型,让语音伪造检测解释更可信

XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models

论文配图:XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
图 1 · 摘自论文原文
  • 用XAI信号引导多模态大模型生成可解释的检测理由
  • 在PartialSpoof数据集上提升内部准确率45%以上
  • 无需训练,适合需要可信解释的AI安全场景

语音深度伪造检测(SDD)系统需提供可信解释以保障决策可靠性。现有解释方法分为两类:传统可解释人工智能(XAI)如基于梯度的归因,生成与模型决策紧密耦合的低级信号,人类难以理解;而基于大语言模型(LLM)的解释常因缺乏启发式证据和任务特定监督,产生泛化且无根据的描述,根源在于SDD领域缺乏高质量的可解释数据集。为此,我们提出一种无需训练的解释框架,将XAI证据与多模态大模型结合,生成有依据且具体的解释。基于PartialSpoof数据集构建了可解释数据集,实验表明引入XAI的方法使内部准确率提升超过45%,并通过人工评估与忠实性检查验证。

原文摘要 · Abstract (English)

Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explainable AI (XAI), such as gradient-based attribution, produces low-level attribution signals tightly coupled with model decisions, and harder to be understood by human than natural language explanations. Meanwhile, large language model (LLM)-based explanation generation often produces generic and ungrounded descriptions due to the lack of heuristic evidence and task-specific supervision, stemming from limited grounded explanation datasets for SDD. We therefore propose a training-free explanation framework that integrates XAI evidence with multimodal LLMs to generate grounded and specific explanations. Using the PartialSpoof dataset, we construct a grounded explanation dataset and show that methods with XAI increase inside accuracy by over 45\%, verified through human evaluation and faithfulness checks.

语音检测可解释AI大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。