arXiv:2511.22471cs.CV2025-11被引 11

无需训练,用DINOv3识别跨生成模型的伪造图像

Rethinking Cross-Generator Image Forgery Detection through DINOv3

  • 利用DINOv3的全局低频结构特征,识别通用伪造线索
  • 仅通过筛选关键特征令牌,检测准确率全面提升
  • 适合需要高效、可解释伪造检测的开发者和研究者

随着生成模型日益多样化和强大,跨生成器检测成为新挑战。现有方法常记忆特定生成模型的痕迹,难以泛化到未知生成器。本研究发现,冻结的视觉基础模型(尤其是DINOv3)在未经微调的情况下已具备强跨生成器检测能力。通过频率、空间与令牌三个视角的系统分析,我们观察到DINOv3依赖全局、低频结构作为弱但可迁移的真伪线索,而非高频的生成器特异性伪影。基于此,我们提出一种无需训练的令牌排序策略,结合轻量线性探测器,筛选出少量与真伪相关的特征令牌。该方法在所有评估数据集上均显著提升检测精度。本研究为理解基础模型在多样化生成器间的泛化机制提供了实证支持,并构建了一个通用、高效且可解释的伪造图像检测基准。

原文摘要 · Abstract (English)

As generative models become increasingly diverse and powerful, cross-generator detection has emerged as a new challenge. Existing detection methods often memorize artifacts of specific generative models rather than learning transferable cues, leading to substantial failures on unseen generators. Surprisingly, this work finds that frozen visual foundation models, especially DINOv3, already exhibit strong cross-generator detection capability without any fine-tuning. Through systematic studies on frequency, spatial, and token perspectives, we observe that DINOv3 tends to rely on global, low-frequency structures as weak but transferable authenticity cues instead of high-frequency, generator-specific artifacts. Motivated by this insight, we introduce a simple, training-free token-ranking strategy followed by a lightweight linear probe to select a small subset of authenticity-relevant tokens. This token subset consistently improves detection accuracy across all evaluated datasets. Our study provides empirical evidence and a feasible hypothesis for understanding why foundation models generalize across diverse generators, offering a universal, efficient, and interpretable baseline for image forgery detection.

伪造检测DINOv3跨生成器无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。