针对AI生成全景图的质量评估与失真敏感显著性预测,提出新模型与优化流程。
Quality Assessment and Distortion-aware Saliency Prediction for AI-Generated Omnidirectional Images
- 基于BLIP-2构建共享编码器模型,评估视觉体验并预测失真敏感区域。
- 在OHF2024数据集上达到当前最优性能,显著提升全景图质量评估精度。
- 适用于VR/AR场景下AI生成全景图的自动优化,推动沉浸式内容发展。
随着人工智能生成内容(AIGC)技术快速发展,AI生成全景图(AIGODI)在虚拟现实(VR)和增强现实(AR)中展现出巨大潜力。然而,这类图像存在独特质量问题,相关质量评估与优化研究仍不足。为此,本文首次研究AIGODI的质量评估与失真感知显著性预测问题,并提出相应优化流程。首先构建包含主观质量评分(三视角)与失真感知显著区域的综合性数据库OHF2024。基于该数据集,提出两个基于BLIP-2共享编码器的模型:用于评估人类视觉体验的BLIP2OIQA,以及用于预测失真敏感显著性的BLIP2OISal。实验表明,两者在人眼视觉体验评价与失真感知显著性预测任务中均达到最先进水平,且可有效用于图像自动优化。相关数据集与代码已开源至https://github.com/IntMeGroup/AIGCOIQA,以促进后续研究。
原文摘要 · Abstract (English)
With the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated omnidirectional images (AIGODIs) hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications. AI generated omnidirectional images exhibit unique quality issues, however, research on the quality assessment and optimization of AI-generated omnidirectional images is still lacking. To this end, this work first studies the quality assessment and distortion-aware saliency prediction problems for AIGODIs, and further presents a corresponding optimization process. Specifically, we first establish a comprehensive database to reflect human feedback for AI-generated omnidirectionals, termed OHF2024, which includes both subjective quality ratings evaluated from three perspectives and distortion-aware salient regions. Based on the constructed OHF2024 database, we propose two models with shared encoders based on the BLIP-2 model to evaluate the human visual experience and predict distortion-aware saliency for AI-generated omnidirectional images, which are named as BLIP2OIQA and BLIP2OISal, respectively. Finally, based on the proposed models, we present an automatic optimization process that utilizes the predicted visual experience scores and distortion regions to further enhance the visual quality of an AI-generated omnidirectional image. Extensive experiments show that our BLIP2OIQA model and BLIP2OISal model achieve state-of-the-art (SOTA) results in the human visual experience evaluation task and the distortion-aware saliency prediction task for AI generated omnidirectional images, and can be effectively used in the optimization process. The database and codes will be released on https://github.com/IntMeGroup/AIGCOIQA to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。