arXiv:2509.08583cs.CV2025-09

针对高分辨率扩散伪造检测,提出轻量级高效定位模型。

EfficientIML: Efficient High-Resolution Image Manipulation Localization

  • 用三阶段轻量级EfficientRWKV捕捉全局与局部特征
  • 在1200+张高分辨率扩散伪造图上实现更优定位精度与速度
  • 适合实时图像取证场景,兼顾性能与计算效率

随着成像设备分辨率不断提升及基于扩散模型的伪造方法兴起,现有仅在传统数据集(拼接、复制移动、物体移除)上训练的检测器缺乏对新型伪造类型的识别能力。为此,我们构建了一个包含1200+张扩散生成伪造图像的高分辨率SIF数据集,并提供语义提取掩码。然而,这给现有方法带来计算资源瓶颈,因其计算复杂度过高。为此,我们提出EfficientIML模型,采用轻量级三阶段EfficientRWKV主干网络,该网络融合状态空间与注意力机制,可并行捕获全局上下文与局部细节;同时引入多尺度监督策略,确保分层预测的一致性。在自建数据集与标准基准上的大量实验表明,该方法在定位性能、FLOPs与推理速度上均优于基于ViT及其他SOTA轻量级基线,验证了其在实时图像取证中的适用性。

原文摘要 · Abstract (English)

With imaging devices delivering ever-higher resolutions and the emerging diffusion-based forgery methods, current detectors trained only on traditional datasets (with splicing, copy-moving and object removal forgeries) lack exposure to this new manipulation type. To address this, we propose a novel high-resolution SIF dataset of 1200+ diffusion-generated manipulations with semantically extracted masks. However, this also imposes a challenge on existing methods, as they face significant computational resource constraints due to their prohibitive computational complexities. Therefore, we propose a novel EfficientIML model with a lightweight, three-stage EfficientRWKV backbone. EfficientRWKV's hybrid state-space and attention network captures global context and local details in parallel, while a multi-scale supervision strategy enforces consistency across hierarchical predictions. Extensive evaluations on our dataset and standard benchmarks demonstrate that our approach outperforms ViT-based and other SOTA lightweight baselines in localization performance, FLOPs and inference speed, underscoring its suitability for real-time forensic applications.

图像伪造检测扩散模型轻量化模型高分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。