arXiv:2506.17969cs.CV2025-06中稿 · ICME 2025被引 5

用CLIP从低层失真到高层语义逐步评估图像质量,更贴近人眼感知。

BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIP

  • 自下而上设计跨尺度注意力,捕捉失真对语义的影响
  • 引入40个质量形容词增强文本表征,连接图像质量与人类语言
  • 在多种基准上表现更优且更鲁棒,适合高质量图像评估场景

图像质量评估(IQA)旨在基于人类主观感知评价图像的感知质量。现有方法通常融合多尺度特征以获得高性能,但多数依赖简单的线性融合,难以充分捕捉失真对语义内容的影响。为此,我们提出一种基于对比语言-图像预训练模型(CLIP)的自下而上图像质量评估方法,命名为BPCLIP,逐步提取低层失真对高层语义的影响。具体地,利用编码器提取输入图像的多尺度特征,并引入自下而上的多尺度交叉注意力模块,以捕捉浅层与深层特征间的关联。此外,通过整合6个维度共40个图像质量形容词,使预训练的CLIP文本编码器生成图像内在质量的表示,从而强化图像质量感知与人类语言之间的联系。该方法在大多数公开的全参考(FR)和无参考(NR)IQA基准上取得更优结果,同时表现出更强的鲁棒性。

原文摘要 · Abstract (English)

Image Quality Assessment (IQA) aims to evaluate the perceptual quality of images based on human subjective perception. Existing methods generally combine multiscale features to achieve high performance, but most rely on straightforward linear fusion of these features, which may not adequately capture the impact of distortions on semantic content. To address this, we propose a bottom-up image quality assessment approach based on the Contrastive Language-Image Pre-training (CLIP, a recently proposed model that aligns images and text in a shared feature space), named BPCLIP, which progressively extracts the impact of low-level distortions on high-level semantics. Specifically, we utilize an encoder to extract multiscale features from the input image and introduce a bottom-up multiscale cross attention module designed to capture the relationships between shallow and deep features. In addition, by incorporating 40 image quality adjectives across six distinct dimensions, we enable the pre-trained CLIP text encoder to generate representations of the intrinsic quality of the image, thereby strengthening the connection between image quality perception and human language. Our method achieves superior results on most public Full-Reference (FR) and No-Reference (NR) IQA benchmarks, while demonstrating greater robustness.

图像质量评估CLIP多尺度特征语义感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。