揭示深度伪造检测器的隐含语义,让黑箱模型可解释
Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors

- 用编码-解码方向对分析检测器内部语义空间
- 发现训练中隐含的真实与虚假特征模式
- 适合需要可信解释的法律、媒体审核场景
深度伪造技术严重威胁信息真实性,亟需可靠的检测方法。现有检测器多仅输出真假二分类结果,缺乏实际应用所需的解释能力。可解释性检测虽已出现,但常依赖人工标注或生成无依据的表面解释。本文采用后处理可解释AI(XAI)技术,通过编码-解码方向对(EDDP)分析先进黑箱检测器的决策过程,揭示其在训练中隐式学习到的真实与虚假特征,实现全局模型理解、空间感知的特征定位及反事实分析,为深度伪造检测提供前所未有的细粒度解释能力。
原文摘要 · Abstract (English)
Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only consider predicting a binary label for whether the input is real or fake, lacking the justification required for real-world applications like legal proceedings. Explainable DF Detection has emerged to address this limitation, but existing techniques frequently fall short by either relying on human annotations for precise artifact localization or generating superficially plausible textual explanations without grounding. This work investigates the use of post-hoc explainable AI (XAI) to analyze the decision-making process of state-of-the-art black-box DF detectors. Specifically, we employ Encoding-Decoding Direction Pairs (EDDP), a technique suitable for uncovering the concept space of DF detectors (their semantic vocabulary) as well as the mechanism for writing and reading concept information to and from internal representations. Our analysis reveals previously hidden real and fake features learned implicitly during detector training, offering nuanced explanations unattainable through conventional methods. This enables global model understanding, spatially aware concept localization, and counterfactual what-if analysis, all contributing to a deeper comprehension of DF detection strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。