arXiv:2410.06044cs.CV2024-10被引 4

用混合超LoRA提升跨模型假图检测能力

HyperDet: Generalizable Detection of Synthesized Images by Generating and Merging A Mixture of Hyper LoRAs

  • 分组SRM滤波器+超网络生成多类LoRA权重
  • 在UnivFD和Fake2M上达最新最好性能
  • 适合需要泛化检测合成图像的研究者

生成视觉模型的兴起使得合成图像日益逼真,亟需有效区分真实与合成图像。现有方法难以准确识别不同生成模型产生的合成图像。本文提出HyperDet框架,通过整合功能各异且轻量的专家检测器共享知识,实现通用检测。HyperDet利用预训练视觉模型提取通用特征,同时捕捉并增强任务特定特征:首先将SRM滤波器分为五组,以高效捕获不同层级的像素伪影;接着用超网络生成具有不同嵌入参数的LoRA权重;最后融合多个LoRA网络形成高效集成模型。此外,提出新目标函数,有效平衡像素与语义伪影的检测。在UnivFD和Fake2M数据集上的大量实验表明,该方法性能领先。本工作为基于预训练大视觉模型构建可泛化的领域特定假图检测器提供了新路径。

原文摘要 · Abstract (English)

The emergence of diverse generative vision models has recently enabled the synthesis of visually realistic images, underscoring the critical need for effectively detecting these generated images from real photos. Despite advances in this field, existing detection approaches often struggle to accurately identify synthesized images generated by different generative models. In this work, we introduce a novel and generalizable detection framework termed HyperDet, which innovatively captures and integrates shared knowledge from a collection of functionally distinct and lightweight expert detectors. HyperDet leverages a large pretrained vision model to extract general detection features while simultaneously capturing and enhancing task-specific features. To achieve this, HyperDet first groups SRM filters into five distinct groups to efficiently capture varying levels of pixel artifacts based on their different functionality and complexity. Then, HyperDet utilizes a hypernetwork to generate LoRA model weights with distinct embedding parameters. Finally, we merge the LoRA networks to form an efficient model ensemble. Also, we propose a novel objective function that balances the pixel and semantic artifacts effectively. Extensive experiments on the UnivFD and Fake2M datasets demonstrate the effectiveness of our approach, achieving state-of-the-art performance. Moreover, our work paves a new way to establish generalizable domain-specific fake image detectors based on pretrained large vision models.

图像检测生成模型伪影分析LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。