用潜在像素一致性检测深度伪造,提升识别精度与解释性。
LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

- 结合图像重建残差与视觉特征进行多模态推理
- 在通用伪造检测上达到领先性能,定位准确率超基准
- 适合关注生成模型漏洞分析与可解释检测的研究者
当前生成模型产生的图像视觉瑕疵较少,削弱了仅依赖表面特征的检测器和解释能力。本文提出LaP-Forensics,一种融合RGB语义与基于重建的取证证据的多模态框架。采用冻结的Stable Diffusion DDIM反演-重建模型生成固定重建参考,其残差图衡量局部一致性。独立投影器编码RGB图像与残差图,再通过结构化Where-What-Why模型预测文本分析与伪影掩码。经监督微调后,使用组相对策略优化(GRPO),奖励函数结合掩码重叠、输出结构与证据参考项。文本侧项鼓励模型引用一致性图,但不验证自由文本真实性。另设图像级头融合RGB与DDIM残差特征。实验显示在UniversalFakeDetect上实现跨生成器检测,在官方SynthScars基准上表现优异。控制变量、反事实等分析证实残差流的有效性,但后处理下的自由文本真实性和可靠性仍待解决。
原文摘要 · Abstract (English)
Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance. We present LaP-Forensics, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence. A frozen Stable Diffusion DDIM inversion-reconstruction model provides a fixed reconstruction reference, and its residual map measures local compatibility with that reference. Independent projectors encode the RGB image and residual map before a structured Where-What-Why model predicts a textual analysis and an artifact mask.Supervised fine-tuning is followed by Group Relative Policy Optimization (GRPO), whose reward combines mask overlap with output-structure and evidence-reference terms. These text-side terms encourage the model to refer to the consistency map but do not constitute a verifier of free-form textual truth. A separate image-level head fuses RGB and DDIM-residual class features. Experiments show cross-generator detection on UniversalFakeDetect and competitive artifact localization on the official SynthScars benchmark. Controlled cue-construction, inversion-horizon, component, reward-term, and counterfactual analyses support the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。