arXiv:2601.08873cs.CVcs.AI2026-01被引 1

用分层注意力模型检测跨域图像伪造,准确率超86%。

ForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection

  • 分三层分析:像素级痕迹、边界异常、语义逻辑,用交叉注意力融合
  • 跨7类数据集平均准确率达86.8%,压缩后仍保持83%精度
  • 可定位伪造区域,结果符合人工专家判断,适合未知伪造场景

AI生成图像和高级编辑工具的普及使传统取证方法在跨域检测中失效。我们提出ForensicFormer,一种分层多尺度框架,通过交叉注意力变换器统一低层痕迹检测、中层边界分析和高层语义推理。相比以往单一范式方法在分布外数据集上低于75%的准确率,本方法在七个多样化测试集(涵盖传统篡改、GAN生成图像及扩散模型输出)上保持86.8%平均准确率,显著优于现有通用检测器。在JPEG压缩(Q=70)下仍达83%准确率(基线仅66%),并实现像素级伪造定位,F1-score为0.76。大量消融实验表明,每一层组件贡献4-10%准确率提升;定性分析显示其提取的特征与人工专家推理一致。本工作融合经典图像取证与现代深度学习,为未知伪造技术的真实场景部署提供可行方案。

原文摘要 · Abstract (English)

The proliferation of AI-generated imagery and sophisticated editing tools has rendered traditional forensic methods ineffective for cross-domain forgery detection. We present ForensicFormer, a hierarchical multi-scale framework that unifies low-level artifact detection, mid-level boundary analysis, and high-level semantic reasoning via cross-attention transformers. Unlike prior single-paradigm approaches, which achieve <75% accuracy on out-of-distribution datasets, our method maintains 86.8% average accuracy across seven diverse test sets, spanning traditional manipulations, GAN-generated images, and diffusion model outputs - a significant improvement over state-of-the-art universal detectors. We demonstrate superior robustness to JPEG compression (83% accuracy at Q=70 vs. 66% for baselines) and provide pixel-level forgery localization with a 0.76 F1-score. Extensive ablation studies validate that each hierarchical component contributes 4-10% accuracy improvement, and qualitative analysis reveals interpretable forensic features aligned with human expert reasoning. Our work bridges classical image forensics and modern deep learning, offering a practical solution for real-world deployment where manipulation techniques are unknown a priori.

图像伪造检测分层推理Transformer跨域检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。