arXiv:2512.21670cs.CVcs.LG2025-12

通过稀疏特征与流形分析,揭示深度伪造检测模型的决策机制。

The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds

  • 用稀疏自编码器分析模型内部表示,定位关键特征。
  • 发现仅少数隐层特征被激活,且流形几何随伪造类型变化。
  • 为构建可解释、鲁棒的检测模型提供新思路,适合安全与可信AI研究者。

深度伪造检测模型虽已实现高准确率,但其决策过程仍不透明。本文提出一种针对视觉-语言模型的机制可解释性框架,结合稀疏自编码器(SAE)对网络内部表示的分析,以及一种新型的取证流形分析,探测模型特征在受控伪造痕迹扰动下的响应。结果表明,每一层仅有少量潜在特征被实际使用,且模型特征流形的几何特性(如内在维度、曲率、特征选择性)会随不同类型的深度伪造痕迹系统性变化。这些发现首次打开了深度伪造检测器的‘黑箱’,使我们能识别对应特定取证痕迹的学得特征,并指导更可解释、更鲁棒模型的开发。

原文摘要 · Abstract (English)

Deepfake detection models have achieved high accuracy in identifying synthetic media, but their decision processes remain largely opaque. In this paper we present a mechanistic interpretability framework for deepfake detection applied to a vision-language model. Our approach combines a sparse autoencoder (SAE) analysis of internal network representations with a novel forensic manifold analysis that probes how the model's features respond to controlled forensic artifact manipulations. We demonstrate that only a small fraction of latent features are actively used in each layer, and that the geometric properties of the model's feature manifold, including intrinsic dimensionality, curvature, and feature selectivity, vary systematically with different types of deepfake artifacts. These insights provide a first step toward opening the "black box" of deepfake detectors, allowing us to identify which learned features correspond to specific forensic artifacts and to guide the development of more interpretable and robust models.

可解释性深度伪造神经流形特征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。