arXiv:2510.04225cs.CVcs.AI2025-10被引 3

通过定位可疑区域再分析,提升对高质伪造图像的检测与解释能力。

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

  • 先定位可疑区域,再结合局部与全局图像重新判断真伪。
  • 在2万张图像数据集上达到领先准确率,且解释更贴近视觉证据。
  • 适合需要可解释性判别的数字取证场景,如媒体审核与内容验证。

AI生成图像的迅猛发展模糊了真实与合成内容的界限,威胁数字真实性。视觉语言模型(VLM)虽能提供自然语言解释,但传统单步分类器常忽略高质量合成图像中的细微痕迹,且缺乏像素级定位。本文提出定位-再审视(Locate-Then-Examine, LTE)框架,采用两阶段VLM方法:首先定位可疑区域,再将这些局部区域与完整图像联合复审,以优化真伪判断及解释。LTE通过区域提议显式关联每项决策与具体视觉证据,实现区域感知推理。为支持训练与评估,我们构建了TRACE数据集,包含20,000张真实与高质量合成图像,配有区域级标注和自动生成的法证解释,由VLM管道生成并经一致性检查与质量控制。在TRACE及多个外部基准上,LTE表现优异,具备更强鲁棒性,并输出人类可理解、区域可信的解释,适用于实际法证部署。

原文摘要 · Abstract (English)

The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs) can provide natural language explanations, but standard one-pass classifiers often miss subtle artifacts in high-quality synthetic images and offer limited grounding in the pixels. We propose Locate-Then-Examine (LTE), a two-stage VLM-based forensic framework that first localizes suspicious regions and then re-examines these crops together with the full image to refine the real vs. AI-generated verdict and its explanation. LTE explicitly links each decision to localized visual evidence through region proposals and region-aware reasoning. To support training and evaluation, we introduce TRACE, a dataset of 20,000 real and high-quality synthetic images with region-level annotations and automatically generated forensic explanations, constructed by a VLM-based pipeline with additional consistency checks and quality control. Across TRACE and multiple external benchmarks, LTE achieves competitive accuracy and improved robustness while providing human-understandable, region-grounded explanations suitable for forensic deployment.

图像伪造检测视觉语言模型可解释性法证分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。