提出首个针对局部感知伪影的检测框架与数据集,提升图像质量评估精度。
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts

- 基于CLIP构建新框架,通过语义关联增强对局部伪影的识别能力
- 在3520张图像上实现显著优于现有方法的检测性能
- 适合图像质量评估、数字内容可信度验证等场景使用
当前图像质量评估方法多关注全局失真(如噪声、模糊),忽视了鬼影、镜头光晕、摩尔纹等局部感知伪影。尽管伪影去除已取得进展,但自动检测仍缺乏系统研究。本文首次正式定义图像感知伪影检测(IPAD)任务,构建包含3,520张图像的数据集(520张实拍、3,000张合成),每张配以像素级掩码,覆盖三类典型伪影。核心挑战在于伪影局部、细微且语义弱,易被漏检。为此,提出IPAD-CLIP框架,在文本与视觉空间增强伪影区分能力,同时保持泛化性。关键思路是:局部伪影常与特定语义上下文强相关,因此学习伪影感知文本嵌入,显式建模物体-伪影关系,使模型能清晰区分干净与含伪影提示。这些文本嵌入作为锚点,引导视觉编码器注意力从高层语义转向低层细微伪影。大量实验表明,IPAD-CLIP在资源高效的前提下显著优于先进异常检测与篡改检测方法。据我们所知,这是首个在数据集与模型层面同时解决多类局部感知伪影检测的研究。
原文摘要 · Abstract (English)
Current image quality assessment methods are heavily biased towards global distortions (e.g., noise, blur), neglecting local perceptual artifacts such as ghosting, lens flare, and moire effects. Although significant progress has been made in artifact removal, the fundamental problem of automatic artifact detection remains largely unexplored. In this paper, we formalize the Image Perceptual Artifact Detection (IPAD) task to address this gap. We contribute a benchmark dataset comprising 3,520 artifact images, including 520 real-captured and 3,000 synthetic samples, each paired with pixel-level masks across three representative artifact categories. The core challenge of IPAD lies in the localized, subtle, and semantically weak nature of these artifacts, which makes them prone to missed detection. To overcome this, we introduce IPAD-CLIP, a novel framework built upon CLIP that enhances artifact discrimination in both textual and visual spaces while preserving generalization capabilities. Our key insight is that local artifacts often exhibit strong correlations with specific semantic contexts. Accordingly, we learn artifact-aware text embeddings to explicitly model the object-artifact relationships, resulting in enhanced representations that clear differentiate between clean and artifact prompts. These text embeddings are then used as anchors to shift the visual encoder's attention from high-level semantics to subtle, low-level artifacts. Extensive experiments demonstrate that IPAD-CLIP offers a resource-efficient adaptation of CLIP for detection, significantly outperforming advanced image anomaly detection and manipulation detection methods on our benchmark. To the best of our knowledge, this is the first study addressing multi-class local perceptual artifact detection in terms of both dataset and model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。