同时预测图像压缩后人机满意率,提升视觉质量优化效率。
Predicting Satisfied User and Machine Ratio for Compressed Images: A Unified Approach
- 用深度学习联合建模人眼与机器对压缩图像的满意比例。
- 在真实标注数据上预训练特征提取器,结合多层特征融合预测。
- 引入残差差异特征学习与注意力聚合,显著提升预测精度。
当前,人类追求高质量图像以获得更好观看体验,机器则依赖高质图像实现更准确的视觉分析。然而图像在消费前常被压缩,导致质量下降。因此,预测压缩图像对人和机器的感知质量具有重要意义,可指导压缩优化策略。本文提出一种统一方法,构建深度学习模型同时预测满足用户比率(SUR)与满足机器比率(SMR)。首先,在大规模带有SMR标注的数据集上预训练特征提取网络,利用多种图像质量模型生成的人类感知相关标签模拟SUR标签获取过程;随后,设计基于MLP-Mixer的网络,通过融合多层特征预测SUR与SMR。提出差异特征残差学习(DFRL)模块以捕捉更具判别性的差异特征,并引入多头注意力聚合与池化(MHAAP)层聚合差异特征并降低冗余。实验表明,所提模型显著优于现有SUR与SMR预测方法。此外,人机感知质量联合学习机制有效提升了两项任务的性能。
原文摘要 · Abstract (English)
Nowadays, high-quality images are pursued by both humans for better viewing experience and by machines for more accurate visual analysis. However, images are usually compressed before being consumed, decreasing their quality. It is meaningful to predict the perceptual quality of compressed images for both humans and machines, which guides the optimization for compression. In this paper, we propose a unified approach to address this. Specifically, we create a deep learning-based model to predict Satisfied User Ratio (SUR) and Satisfied Machine Ratio (SMR) of compressed images simultaneously. We first pre-train a feature extractor network on a large-scale SMR-annotated dataset with human perception-related quality labels generated by diverse image quality models, which simulates the acquisition of SUR labels. Then, we propose an MLP-Mixer-based network to predict SUR and SMR by leveraging and fusing the extracted multi-layer features. We introduce a Difference Feature Residual Learning (DFRL) module to learn more discriminative difference features. We further use a Multi-Head Attention Aggregation and Pooling (MHAAP) layer to aggregate difference features and reduce their redundancy. Experimental results indicate that the proposed model significantly outperforms state-of-the-art SUR and SMR prediction methods. Moreover, our joint learning scheme of human and machine perceptual quality prediction tasks is effective at improving the performance of both.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。