构建首个针对监控图像伪造检测的大型数据集,提升真实场景下伪造定位能力。
SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation

- 用多模态大模型生成语义感知的细粒度伪造图像,覆盖多样监控场景。
- 包含13.7万张不同分辨率和篡改类型的伪造图,显著挑战现有检测器性能。
- 适用于安全取证、视频真实性验证的研究者与实际应用开发者。
我们提出了监控图像伪造检测与定位数据集SurFITR,以应对开源图像生成模型引发的视觉证据伪造风险。现有伪造模型在全图合成或大区域篡改的物体中心图像上训练,难以泛化至监控场景——后者通常为局部、细微的篡改,且存在视角多样、目标小或遮挡、画质低等问题。SurFITR通过多模态大模型驱动的流水线生成大量具有法医价值的图像,支持跨场景的语义感知精细编辑。数据集包含超过13.7万张不同分辨率和篡改类型的真实监控风格伪造图像,使用多种图像编辑模型生成。大量实验表明,现有检测器在SurFITR上性能大幅下降,而基于SurFITR训练可显著提升域内及跨域检测表现。该数据集已公开于GitHub。
原文摘要 · Abstract (English)
We present the Surveillance Forgery Image Test Range (SurFITR), a dataset for surveillance-style image forgery detection and localisation, in response to recent advances in open-access image generation models that raise concerns about falsifying visual evidence. Existing forgery models, trained on datasets with full-image synthesis or large manipulated regions in object-centric images, struggle to generalise to surveillance scenarios. This is because tampering in surveillance imagery is typically localised and subtle, occurring in scenes with varied viewpoints, small or occluded subjects, and lower visual quality. To address this gap, SurFITR provides a large collection of forensically valuable imagery generated via a multimodal LLM-powered pipeline, enabling semantically aware, fine-grained editing across diverse surveillance scenes. It contains over 137k tampered images with varying resolutions and edit types, generated using multiple image editing models. Extensive experiments show that existing detectors degrade significantly on SurFITR, while training on SurFITR yields substantial improvements in both in-domain and cross-domain performance. SurFITR is publicly available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。