arXiv:2503.06641cs.CV2025-03被引 5

提出新方法提升图像复杂度表示的稳定性与准确性。

CLICv2: Image Complexity Representation via Content Invariance Contrastive Learning

  • 通过随机位移图像块生成正样本,避免内容偏差。
  • 在IC9600数据集上皮尔逊相关系数达0.89,斯皮尔曼相关系数0.91。
  • 适合图像质量评估、视觉感知研究等场景使用。

无监督图像复杂度表示常因正样本选择偏差和对图像内容敏感而受限。本文提出CLICv2,一种基于内容不变性对比学习的框架。与CLIC通过裁剪生成正样本(引入偏差)不同,CLICv2采用随机方向位移图像块的方法,在对应位置的块作为正样本对,实现内容不变学习。此外,提出块级对比损失,增强局部复杂度表征并抑制内容干扰。为进一步消除内容干扰,引入掩码图像建模作为辅助任务,但其建模目标为被掩码块的熵,通过未掩码块信息恢复整体图像熵,从而获得全局复杂度感知能力。在IC9600数据集上的大量实验表明,CLICv2显著优于现有无监督方法,在皮尔逊相关系数(PCC)和斯皮尔曼等级相关系数(SRCC)上均表现更优,实现了无正样本偏差的内容不变复杂度表示。

原文摘要 · Abstract (English)

Unsupervised image complexity representation often suffers from bias in positive sample selection and sensitivity to image content. We propose CLICv2, a contrastive learning framework that enforces content invariance for complexity representation. Unlike CLIC, which generates positive samples via cropping-introducing positive pairs bias-our shifted patchify method applies randomized directional shifts to image patches before contrastive learning. Patches at corresponding positions serve as positive pairs, ensuring content-invariant learning. Additionally, we propose patch-wise contrastive loss, which enhances local complexity representation while mitigating content interference. In order to further suppress the interference of image content, we introduce Masked Image Modeling as an auxiliary task, but we set its modeling objective as the entropy of masked patches, which recovers the entropy of the overall image by using the information of the unmasked patches, and then obtains the global complexity perception ability. Extensive experiments on IC9600 demonstrate that CLICv2 significantly outperforms existing unsupervised methods in PCC and SRCC, achieving content-invariant complexity representation without introducing positive pairs bias.

图像复杂度对比学习无监督表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。