新方法融合全局与局部感知,更贴近人眼对图像质量的判断。
Scene Perceived Image Perceptual Score (SPIPS): combining global and local perception for image quality assessment
- 分离高层语义与低层感知特征,分别处理再融合。
- 在多个数据集上优于现有模型,更符合人类主观评分。
- 适合评估AI生成或深度后处理图像的质量。
人工智能的快速发展和智能手机的普及导致真实(相机拍摄)和虚拟(AI生成)图像数据呈指数级增长。这凸显了亟需能够准确反映人类视觉感知的图像质量评估(IQA)方法。传统IQA技术主要依赖空间特征——如信噪比、局部结构失真和纹理不一致——来识别伪影。虽然对未经处理或常规修改的图像有效,但在现代基于深度神经网络(DNN)的图像后处理背景下表现不足。DNN驱动的图像生成、增强和修复模型显著提升了视觉质量,但使准确评估变得更为复杂。为此,我们提出一种新型IQA方法,弥合深度学习与人类感知之间的差距。该模型将深层特征解耦为高层语义信息和低层感知细节,分别处理后与传统IQA指标结合,构建更全面的评估框架。这种混合设计使模型能同时评估全局上下文与精细图像细节,更贴合人类视觉过程(先理解整体结构,再关注细微元素)。最后阶段使用多层感知机(MLP)将整合特征映射为简洁的质量得分。实验表明,该方法在一致性上优于现有IQA模型,更符合人类主观判断。
原文摘要 · Abstract (English)
The rapid advancement of artificial intelligence and widespread use of smartphones have resulted in an exponential growth of image data, both real (camera-captured) and virtual (AI-generated). This surge underscores the critical need for robust image quality assessment (IQA) methods that accurately reflect human visual perception. Traditional IQA techniques primarily rely on spatial features - such as signal-to-noise ratio, local structural distortions, and texture inconsistencies - to identify artifacts. While effective for unprocessed or conventionally altered images, these methods fall short in the context of modern image post-processing powered by deep neural networks (DNNs). The rise of DNN-based models for image generation, enhancement, and restoration has significantly improved visual quality, yet made accurate assessment increasingly complex. To address this, we propose a novel IQA approach that bridges the gap between deep learning methods and human perception. Our model disentangles deep features into high-level semantic information and low-level perceptual details, treating each stream separately. These features are then combined with conventional IQA metrics to provide a more comprehensive evaluation framework. This hybrid design enables the model to assess both global context and intricate image details, better reflecting the human visual process, which first interprets overall structure before attending to fine-grained elements. The final stage employs a multilayer perceptron (MLP) to map the integrated features into a concise quality score. Experimental results demonstrate that our method achieves improved consistency with human perceptual judgments compared to existing IQA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。