用低秩微调提升AI生成图像在真实场景下的检测鲁棒性
Boosting Robust AIGI Detection with LoRA-based Pairwise Training
- 基于视觉基础模型,通过模拟复杂失真和尺寸变化进行微调
- 在NTIRE挑战赛中取得第三名,显著提升真实场景下检测性能
- 适合关注AI内容安全与对抗性干扰防御的研究者
AI生成图像的泛滥催生了实用检测方法的需求。尽管现有AIGI检测器在干净数据集上表现优异,但在真实环境中(存在不可预测的复杂失真)性能常大幅下降。为此,我们提出一种基于LoRA的成对训练(LPT)策略,专为应对严重失真下的鲁棒检测而设计。核心包括:针对视觉基础模型的定向微调、训练阶段对验证与测试集数据分布的刻意模拟,以及独特的成对训练机制。具体地,引入失真与尺寸模拟以更贴近真实分布;依托视觉基础模型的强大表征能力,微调实现检测;通过成对训练解耦泛化性与鲁棒性优化。实验表明,该方法在NTIRE‘In the Wild’ AIGI检测挑战赛中获得第三名。
原文摘要 · Abstract (English)
The proliferation of highly realistic AI-Generated Image (AIGI) has necessitated the development of practical detection methods. While current AIGI detectors perform admirably on clean datasets, their detection performance frequently decreases when deployed "in the wild", where images are subjected to unpredictable, complex distortions. To resolve the critical vulnerability, we propose a novel LoRA-based Pairwise Training (LPT) strategy designed specifically to achieve robust detection for AIGI under severe distortions. The core of our strategy involves the targeted finetuning of a visual foundation model, the deliberate simulation of data distribution during the training phase, and a unique pairwise training process. Specifically, we introduce distortion and size simulations to better fit the distribution from the validation and test sets. Based on the strong visual representation capability of the visual foundation model, we finetune the model to achieve AIGI detection. The pairwise training is utilized to improve the detection via decoupling the generalization and robustness optimization. Experiments show that our approach secured the 3th placement in the NTIRE Robust AI-Generated Image Detection in the Wild challenge
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。