提出新立体视频质量评估方法,可精准识别雾霾等视觉干扰影响。
Subjective and Objective Quality Assessment Methods of Stereoscopic Videos with Visibility Affecting Distortions
- 构建含360个带雾化失真的高清立体视频数据集,通过24人主观测试得评分。
- 基于双目图像的自然场景统计特性,用广义高斯分布建模多尺度特征差异。
- 无需先验知觉信息,对雾霾等干扰有强判别力,适合真实场景视频质量检测。
本文提出两项主要贡献:1)构建了一个全高清分辨率的立体视频数据集,包含12个参考视频和360个受不同等级雾化与霾化影响的失真视频。通过对原始左右视图施加五级雾/霾环境模拟生成测试刺激,并邀请24名受试者进行主观评价,计算出差异平均意见分数(DMOS)作为该数据集的质量表征;2)提出一种不依赖观感与失真先验信息的立体视频质量评估模型(OU+DU)。该方法从立体视频中构建单眼视图(环状帧),将其划分为非重叠块,分析纯净与失真视频各片段的自然场景统计(NSS)特性,使用单变量广义高斯分布(UGGD)对特征进行建模,计算多空间尺度与多方向球面可调金字塔分解下的参数α、η,证明其具有失真可区分性。进一步对纯净与失真视频特征集进行多元高斯(MVG)建模,计算均值向量与协方差矩阵,并通过巴氏距离度量两者间差异,以估计感知偏差。最终融合两类距离度量,输出整体质量评分。该算法在IRCCYN、LFOVIAS3DPh1、LFOVIAS3DPh2及自建的VAD立体视频数据集上验证,表现稳定且优于现有主流2D/3D图像与视频质量评估方法。
原文摘要 · Abstract (English)
We present two major contributions in this work: 1) we create a full HD resolution stereoscopic (S3D) video dataset comprised of 12 reference and 360 distorted videos. The test stimuli are produced by simulating the five levels of fog and haze ambiances on the pristine left and right video sequences. We perform subjective analysis on the created video dataset with 24 viewers and compute Difference Mean Opinion Scores (DMOS) as quality representative of the dataset, 2) an Opinion Unaware (OU) and Distortion Unaware (DU) video quality assessment model is developed for S3D videos. We construct cyclopean frames from the individual views of an S3D video and partition them into nonoverlapping blocks. We analyze the Natural Scene Statistics (NSS) of all patches of pristine and test videos, and empirically model the NSS features with Univariate Generalized Gaussian Distribution (UGGD). We compute UGGD model parameters (α, \b{eta}) at multiple spatial scales and multiple orientations of spherical steerable pyramid decomposition and show that the UGGD parameters are distortion discriminable. Further, we perform Multivariate Gaussian (MVG) modeling on the pristine and distorted video feature sets and compute the corresponding mean vectors and covariance matrices of MVG fits. We compute the Bhattacharyya distance measure between mean vectors and covariance matrices to estimate the perceptual deviation of a test video from pristine video set. Finally, we pool both distance measures to estimate the overall quality score of an S3D video. The performance of the proposed objective algorithm is verified on the popular S3D video datasets such as IRCCYN, LFOVIAS3DPh1, LFOVIAS3DPh2 and the proposed VAD stereo dataset. The algorithm delivers consistent performance across all datasets and shows competitive performance against off-the-shelf 2D and 3D image and video quality assessment algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。