arXiv:2409.07650cs.CV2024-09被引 7

用大模型中间特征提升图像质量评估,无需训练就更准。

Foundation Models Boost Low-Level Perceptual Similarity Metrics

  • 用大模型中间层特征计算图像相似度,不依赖最终输出。
  • 零训练下超越传统和顶尖学习型评估方法。
  • 适合图像质量评估研究者快速部署高精度指标。

在基于深度学习的全参考图像质量评估(FR-IQA)中,通常通过预训练卷积神经网络或近期的Transformer网络提取特征,并以特征距离衡量失真图像与参考图像间的感知相似性。这些中间特征常需额外微调或添加神经网络层,以使最终评分符合人类判断。目前多数基于基础模型的IQA方法仅依赖最后一层或嵌入向量进行评分。本文探索了尚未被充分挖掘的基础模型中间特征在低层感知相似性度量中的潜力。结果表明,中间特征更具有效性;且无需任何训练,仅通过特征间距离即可实现对传统及最先进学习型度量的超越。

原文摘要 · Abstract (English)

For full-reference image quality assessment (FR-IQA) using deep-learning approaches, the perceptual similarity score between a distorted image and a reference image is typically computed as a distance measure between features extracted from a pretrained CNN or more recently, a Transformer network. Often, these intermediate features require further fine-tuning or processing with additional neural network layers to align the final similarity scores with human judgments. So far, most IQA models based on foundation models have primarily relied on the final layer or the embedding for the quality score estimation. In contrast, this work explores the potential of utilizing the intermediate features of these foundation models, which have largely been unexplored so far in the design of low-level perceptual similarity metrics. We demonstrate that the intermediate features are comparatively more effective. Moreover, without requiring any training, these metrics can outperform both traditional and state-of-the-art learned metrics by utilizing distance measures between the features.

图像质量基础模型感知相似性零训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。