arXiv:2606.18738cs.SD2026-06

提出可定位伪造声谱图异常区域并生成可验证解释的框架

GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis

论文配图:GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis
图 1 · 摘自论文原文
  • 通过选择顶部K个异常区域并生成结构化声学解释
  • 在自建数据集上实现更精准的伪造特征定位与解释质量
  • 适合需要可追溯证据的语音鉴伪场景

语音生成技术的进步使合成语音愈发逼真。尽管现代分类模型在深度伪造检测中表现优异,但无法提供类似‘何处出现伪造线索’及‘其声学含义’的证据,限制了其在司法鉴定中的应用。手动分析完整声谱图成本高,需聚焦诊断性强的区域。现有可解释性方法难以将上下文属性与局部证据关联,导致解释难以验证。为此,我们提出GRIDEX:一个给定深度伪造声谱图后,生成结构化司法解释的流水线。该流程(i)选出声谱图中前K个异常区域,(ii)为每个异常生成包含时间、频谱、音位信息和解释文本的分类声学字段说明。据我们所知,这是首个基于区域定位生成结构化司法解释的框架。GRIDEX采用两阶段训练范式,结合监督微调(SFT)与组相对策略优化(GRPO)。在自建数据集上的实验表明,其在伪造特征定位与解释质量上优于强视觉语言模型基线。数据集与代码将在发表后公开。

原文摘要 · Abstract (English)

The advancement of speech generation technologies has made artificial speech increasingly realistic. Although modern classification models can achieve high accuracy when it comes to deepfake detection, they do not produce evidences such as indicating where spoof cues appear in the spectrogram and what they imply acoustically, limiting their usefulness in forensic settings. Manual analysis of full spectrograms is resource-intensive, so evidence should narrow attention to the most diagnostic regions. Moreover, existing explainability methods have limited capabilities in connecting contextual attributes to localized evidence, making explanations harder to verify. To overcome this limitation, we propose GRIDEX, a pipeline that, when given a deepfake spectrogram, generates forensic explanations of its anomalies. The pipeline (i) selects top-K anomalous regions in the spectrogram and (ii) produces an explanation for each anomaly. The explanations follow a schema of categorical acoustic fields, including temporal, spectral, phonetic information and interpretation text. To our knowledge, this is the first framework to generate structured forensic explanations using regional grounding for deepfake spectrograms. GRIDEX is trained with a two-stage learning paradigm that combines supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO). Experiments on our dataset show improved artifact localization and explanation quality over strong vision-language model (VLM) baselines. The dataset and code will be released upon publication.

语音伪造可解释性司法鉴伪声谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。