构建35万张图像伪造定位数据集,提升AI篡改检测精度与泛化能力
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization
- 基于关键点对齐与语义相似性自动标注伪造区域掩码
- 在新数据集上实现62.5%的IoU,领先现有方法5.1个百分点
- 适用于检测未知编辑模型,对图像退化具有强鲁棒性
prompt-based AI图像编辑的普及加剧了恶意内容伪造和虚假信息风险,但针对此类技术的伪造定位方法仍严重不足。为此,我们提出一种全自动掩码标注框架,利用关键点对齐与语义空间相似性生成精确的伪造区域真值掩码。基于此框架,构建了PromptForge-350k数据集,覆盖四种主流prompt-based AI图像编辑模型,缓解该领域数据稀缺问题。同时提出ICL-Net,采用三流主干结构与跨图像对比学习,有效捕捉鲁棒且通用的取证特征。大量实验表明,该方法在PromptForge-350k上达到62.5% IoU,优于现有SOTA方法5.1%;对常见图像退化仅造成小于1%的性能下降,并在未见编辑模型上实现41.5%的平均IoU,展现出良好泛化能力。
原文摘要 · Abstract (English)
The rapid democratization of prompt-based AI image editing has recently exacerbated the risks associated with malicious content fabrication and misinformation. However, forgery localization methods targeting these emerging editing techniques remain significantly under-explored. To bridge this gap, we first introduce a fully automated mask annotating framework that leverages keypoint alignment and semantic space similarity to generate precise ground-truth masks for edited regions. Based on this framework, we construct PromptForge-350k, a large-scale forgery localization dataset covering four state-of-the-art prompt-based AI image editing models, thereby mitigating the data scarcity in this domain. Furthermore, we propose ICL-Net, an effective forgery localization network featuring a triple-stream backbone and intra-image contrastive learning. This design enables the model to capture highly robust and generalizable forensic features. Extensive experiments demonstrate that our method achieves an IoU of 62.5% on PromptForge-350k, outperforming SOTA methods by 5.1%. Additionally, it exhibits strong robustness against common degradations with an IoU drop of less than 1%, and shows promising generalization capabilities on unseen editing models, achieving an average IoU of 41.5%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。