arXiv:2502.14013cs.GRcs.AI2025-02被引 1

评估AI图像放大后的视觉吸引力,发现Real-ESRGAN和BSRGAN表现最佳。

Appeal prediction for AI up-scaled Images

  • 构建包含1496张图像的标注数据集,涵盖5种放大方法。
  • 真实用户评价显示Real-ESRGAN与BSRGAN在视觉吸引力上领先。
  • 提出两种新模型,基于迁移学习的预测效果最优,适合图像质量评估研究者。

由于机器学习的进步,基于深度神经网络(DNN)或AI的图像放大算法日益流行。已有多种使用CNN、GAN或混合方法的放大模型被提出,但多数仅通过PSNR、SSIM或少量样例图评估性能。本文填补了真实世界图像和主观评价的空白,构建了一个包含136张原始图像和五种放大方法(Real-ESRGAN、BSRGAN、waifu2x、KXNet、Lanczos)的数据集,共生成1496张标注图像。图像吸引力标签通过开源工具AVRate Voyager进行众包标注。实验表明,Real-ESRGAN和BSRGAN在主观评价中表现最佳。我们还训练了深度神经网络以识别放大方法,取得了良好效果。此外,评估了现有图像吸引力与质量模型,均表现不佳,因此我们提出了两种新方法:一种基于迁移学习,性能最优;另一种结合信号特征与随机森林,整体表现良好。所有数据与代码均已开源,支持开放科学研究。

原文摘要 · Abstract (English)

DNN- or AI-based up-scaling algorithms are gaining in popularity due to the improvements in machine learning. Various up-scaling models using CNNs, GANs or mixed approaches have been published. The majority of models are evaluated using PSRN and SSIM or only a few example images. However, a performance evaluation with a wide range of real-world images and subjective evaluation is missing, which we tackle in the following paper. For this reason, we describe our developed dataset, which uses 136 base images and five different up-scaling methods, namely Real-ESRGAN, BSRGAN, waifu2x, KXNet, and Lanczos. Overall the dataset consists of 1496 annotated images. The labeling of our dataset focused on image appeal and has been performed using crowd-sourcing employing our open-source tool AVRate Voyager. We evaluate the appeal of the different methods, and the results indicate that Real-ESRGAN and BSRGAN are the best. Furthermore, we train a DNN to detect which up-scaling method has been used, the trained models have a good overall performance in our evaluation. In addition to this, we evaluate state-of-the-art image appeal and quality models, here none of the models showed a high prediction performance, therefore we also trained two own approaches. The first uses transfer learning and has the best performance, and the second model uses signal-based features and a random forest model with good overall performance. We share the data and implementation to allow further research in the context of open science.

图像放大视觉评估数据集AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。