提出新型轻量级光场图像质量评估方法,提升沉浸式媒体体验。
Light Field Image Quality Assessment With Auxiliary Learning Based on Depthwise and Anglewise Separable Convolutions
- 采用空间-角度分离卷积,高效提取光场图像多维特征。
- 引入辅助学习机制,显著降低预测误差42%以上。
- 适合用于低延迟沉浸式媒体传输与质量监控场景。
在多媒体广播中,无参考图像质量评估(NR-IQA)用于指示用户感知的质量体验(QoE),并支持智能数据传输以优化用户体验。本文提出一种改进的无参考光场图像质量评估(NR-LFIQA)指标,适用于未来沉浸式媒体广播服务。首先,将深度可分离卷积(DSC)扩展至光场图像(LFI)的空间域,提出“光场深度可分离卷积”(LF-DSC),高效提取空间特征。其次,进一步理论扩展至光场的视角空间,提出新颖的“光场角度可分离卷积”(LF-ASC),能够以低复杂度同时提取空间与角度特征,实现全面质量评估。第三,将空间与角度特征估计作为辅助任务,为原生NR-LFIQA任务提供空间与角度质量提示。据我们所知,这是首次探索基于空间-角度提示的深度辅助学习在NR-LFIQA中的应用。在主流光场图像数据集Win5-LID和SMART上进行实验,对比主流全参考IQA指标及当前最优的NR-LFIQA方法。结果表明,所提方法在Win5-LID和SMART上相比次优基准指标,预测误差分别降低42.86%和45.95%;在特定失真类型挑战场景下,误差降幅超过60%。
原文摘要 · Abstract (English)
In multimedia broadcasting, no-reference image quality assessment (NR-IQA) is used to indicate the user-perceived quality of experience (QoE) and to support intelligent data transmission while optimizing user experience. This paper proposes an improved no-reference light field image quality assessment (NR-LFIQA) metric for future immersive media broadcasting services. First, we extend the concept of depthwise separable convolution (DSC) to the spatial domain of light field image (LFI) and introduce "light field depthwise separable convolution (LF-DSC)", which can extract the LFI's spatial features efficiently. Second, we further theoretically extend the LF-DSC to the angular space of LFI and introduce the novel concept of "light field anglewise separable convolution (LF-ASC)", which is capable of extracting both the spatial and angular features for comprehensive quality assessment with low complexity. Third, we define the spatial and angular feature estimations as auxiliary tasks in aiding the primary NR-LFIQA task by providing spatial and angular quality features as hints. To the best of our knowledge, this work is the first exploration of deep auxiliary learning with spatial-angular hints on NR-LFIQA. Experiments were conducted in mainstream LFI datasets such as Win5-LID and SMART with comparisons to the mainstream full reference IQA metrics as well as the state-of-the-art NR-LFIQA methods. The experimental results show that the proposed metric yields overall 42.86% and 45.95% smaller prediction errors than the second-best benchmarking metric in Win5-LID and SMART, respectively. In some challenging cases with particular distortion types, the proposed metric can reduce the errors significantly by more than 60%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。