arXiv:2602.09475cs.CVcs.AI2026-02被引 1

用几百个标注样本即可高效检测图像生成中的伪影,显著降低数据成本。

ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs

  • 基于上下文学习与文本指令优化的多组件架构,激活预训练视觉语言模型的检测能力。
  • 在五个主流伪影检测数据集上达到当前最佳性能,仅需数百张标注图像。
  • 适用于多种伪影类型及AIGC检测任务,通用性强且可快速适配新生成器。

现代图像生成器生成的图像极为逼真,仅通过扭曲的手部或变形物体等伪影可识别其合成来源。检测这些伪影至关重要:缺乏检测能力将无法评估生成器性能或训练奖励模型以改进生成质量。现有检测方法需在数万张标注图像上微调视觉语言模型(VLM),但每当生成器演进或出现新型伪影时,重复此过程成本高昂。我们发现,预训练的VLM已隐含检测伪影所需知识——只需合适的结构支撑,每类伪影仅需数百张标注样本即可激活该能力。我们的系统ArtifactLens在五个主流人类伪影基准上取得当前最优表现(首次跨数据集统一评估),且标注数据量级大幅降低。该结构包含多组件架构,融合上下文学习与文本指令优化,并对各模块提出创新改进。方法可泛化至其他伪影类型(如物体形态、动物解剖结构、实体交互)以及AIGC检测任务。

原文摘要 · Abstract (English)

Modern image generators produce strikingly realistic images, where only artifacts like distorted hands or warped objects reveal their synthetic origin. Detecting these artifacts is essential: without detection, we cannot benchmark generators or train reward models to improve them. Current detectors fine-tune VLMs on tens of thousands of labeled images, but this is expensive to repeat whenever generators evolve or new artifact types emerge. We show that pretrained VLMs already encode the knowledge needed to detect artifacts - with the right scaffolding, this capability can be unlocked using only a few hundred labeled examples per artifact category. Our system, ArtifactLens, achieves state-of-the-art on five human artifact benchmarks (the first evaluation across multiple datasets) while requiring orders of magnitude less labeled data. The scaffolding consists of a multi-component architecture with in-context learning and text instruction optimization, with novel improvements to each. Our methods generalize to other artifact types - object morphology, animal anatomy, and entity interactions - and to the distinct task of AIGC detection.

伪影检测VLM少样本AIGC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。