arXiv:2412.00665cs.CV2024-12被引 2

用少量数据训练出能识别多种生成图像的通用检测器。

Learning on Less: Constraining Pre-trained Model Learning for Generalizable Diffusion-Generated Image Detection

  • 通过随机掩码约束模型学习特定生成模型特征。
  • 仅用1%数据即超越现有最佳方法,平均准确率提升13.6%。
  • 适合需要高效泛化能力的图像真实性检测场景。

扩散模型可生成高度逼真的图像,带来虚假信息风险并削弱公众信任。当前方法在检测未见过的扩散模型生成图像时泛化能力不足。我们重新评估了在大规模真实图像上预训练模型的有效性,发现:1)预训练模型能有效聚类真实图像特征;2)带有预训练权重的模型可在特定训练步数逼近最优泛化解,但极不稳定。基于此,我们提出简单有效的训练方法 Learning on Less(LoL)。LoL采用随机掩码机制,限制模型对某一类扩散模型特有模式的学习,使其聚焦于更通用的图像内容。这利用了预训练权重的内在优势,实现更稳定的最优泛化,从而提取出区分各类生成图像与真实图像的通用特征。在GenImage基准上的大量实验表明,仅使用1%训练数据,LoL显著优于当前最先进方法,在八种不同模型生成图像上的平均准确率提升13.6%。

原文摘要 · Abstract (English)

Diffusion Models enable realistic image generation, raising the risk of misinformation and eroding public trust. Currently, detecting images generated by unseen diffusion models remains challenging due to the limited generalization capabilities of existing methods. To address this issue, we rethink the effectiveness of pre-trained models trained on large-scale, real-world images. Our findings indicate that: 1) Pre-trained models can cluster the features of real images effectively. 2) Models with pre-trained weights can approximate an optimal generalization solution at a specific training step, but it is extremely unstable. Based on these facts, we propose a simple yet effective training method called Learning on Less (LoL). LoL utilizes a random masking mechanism to constrain the model's learning of the unique patterns specific to a certain type of diffusion model, allowing it to focus on less image content. This leverages the inherent strengths of pre-trained weights while enabling a more stable approach to optimal generalization, which results in the extraction of a universal feature that differentiates various diffusion-generated images from real images. Extensive experiments on the GenImage benchmark demonstrate the remarkable generalization capability of our proposed LoL. With just 1% training data, LoL significantly outperforms the current state-of-the-art, achieving a 13.6% improvement in average ACC across images generated by eight different models.

图像检测扩散模型小样本学习通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。