无监督学习图像复杂度表示,提升视觉模型性能。
CLIC: Contrastive Learning Framework for Unsupervised Image Complexity Representation
- 基于对比学习设计正负样本策略,提取复杂度感知特征。
- 仅需少量标注数据微调,性能媲美有监督方法。
- 避免主观偏差,适合复杂度敏感的下游任务。
图像复杂度作为基础视觉属性,显著影响人类感知和计算机视觉模型表现,但准确评估仍具挑战。传统指标如信息熵和压缩比常给出粗糙不可靠的估计;数据驱动方法依赖昂贵的人工标注,且受主观偏见影响。为此,我们提出无监督框架CLIC,基于对比学习学习图像复杂度表征。CLIC从无标签数据中学习复杂度感知特征,无需耗时标注。我们设计新颖的正负样本选择策略以增强特征区分性,并引入基于图像先验的复杂度感知损失函数,进一步约束学习过程。大量实验验证了CLIC在捕捉图像复杂度上的有效性。在仅用少量IC9600标注样本微调后,性能可媲美有监督方法。将CLIC应用于下游任务时,均能持续提升性能。值得注意的是,CLIC的预训练与应用全过程均无主观偏见。
原文摘要 · Abstract (English)
As a fundamental visual attribute, image complexity significantly influences both human perception and the performance of computer vision models. However, accurately assessing and quantifying image complexity remains a challenging task. (1) Traditional metrics such as information entropy and compression ratio often yield coarse and unreliable estimates. (2) Data-driven methods require expensive manual annotations and are inevitably affected by human subjective biases. To address these issues, we propose CLIC, an unsupervised framework based on Contrastive Learning for learning Image Complexity representations. CLIC learns complexity-aware features from unlabeled data, thereby eliminating the need for costly labeling. Specifically, we design a novel positive and negative sample selection strategy to enhance the discrimination of complexity features. Additionally, we introduce a complexity-aware loss function guided by image priors to further constrain the learning process. Extensive experiments validate the effectiveness of CLIC in capturing image complexity. When fine-tuned with a small number of labeled samples from IC9600, CLIC achieves performance competitive with supervised methods. Moreover, applying CLIC to downstream tasks consistently improves performance. Notably, both the pretraining and application processes of CLIC are free from subjective bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。