构建1万张全景图像质量数据集,提出可生成质量描述的新型评估模型。
Omnidirectional Image Quality Captioning: A Large-scale Database and A New Model
- 构建包含10,000张全景图的OIQ-10K数据库,涵盖均匀与非均匀失真。
- 提出IQCaption360模型,能生成基于文本模板的质量描述,性能显著优于现有方法。
- 首次融合人类视觉行为数据,支持复杂失真场景下的图像质量评估,适合全景内容开发者使用。
全景图像应用迅速发展,亟需有效的全景图像质量评估(OIQA)方法。现有方法多在均匀失真图像上训练,难以迁移至非均匀失真场景。本文开展迄今最大的全景图像质量研究,构建包含10,000张全景图像的OIQ-10K数据库,涵盖均匀与非均匀失真,并通过全面的心理物理学实验收集人类主观评分、失真空间分布(局部或全局)以及被试的头部和眼动轨迹。此外,提出一种新型多任务自适应特征调制的OIQA模型IQCaption360,可生成基于文本模板的质量描述。大量实验表明,IQCaption360在新提出的OIQ-10K数据集上显著优于现有先进方法。相关数据集及源码已公开于https://github.com/WenJuing/IQCaption360。
原文摘要 · Abstract (English)
The fast growing application of omnidirectional images calls for effective approaches for omnidirectional image quality assessment (OIQA). Existing OIQA methods have been developed and tested on homogeneously distorted omnidirectional images, but it is hard to transfer their success directly to the heterogeneously distorted omnidirectional images. In this paper, we conduct the largest study so far on OIQA, where we establish a large-scale database called OIQ-10K containing 10,000 omnidirectional images with both homogeneous and heterogeneous distortions. A comprehensive psychophysical study is elaborated to collect human opinions for each omnidirectional image, together with the spatial distributions (within local regions or globally) of distortions, and the head and eye movements of the subjects. Furthermore, we propose a novel multitask-derived adaptive feature-tailoring OIQA model named IQCaption360, which is capable of generating a quality caption for an omnidirectional image in a manner of textual template. Extensive experiments demonstrate the effectiveness of IQCaption360, which outperforms state-of-the-art methods by a significant margin on the proposed OIQ-10K database. The OIQ-10K database and the related source codes are available at https://github.com/WenJuing/IQCaption360.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。