arXiv:2412.06003cs.CVeess.IV2024-12

用知识蒸馏提升AR图像质量评估的表征能力,解决数据少、失真难测的问题。

Enhancing Content Representation for AR Image Quality Assessment Using Knowledge Distillation

  • 通过自监督视觉变压器蒸馏,增强失真图像的特征表示
  • 在ARIQA数据集上,三种模型均超越现有方法,最高提升0.12分
  • 适合做AR质量评估、图像重建与感知建模的研究者参考

增强现实(AR)通过将数字内容叠加到真实环境,广泛应用于娱乐、教育、医疗和工业培训等领域。然而,视觉混淆和传统失真常引发用户不适,因此评估AR体验质量至关重要。由于数据稀缺及AR特性独特,构建有效质量评估指标仍具挑战。本文提出一种基于深度学习的客观评估方法,包含四步:(1)微调自监督预训练视觉变换器提取参考图显著特征,并蒸馏至失真图像以增强表征;(2)通过计算位移表示量化失真;(3)使用交叉注意力解码器捕捉感知质量特征;(4)引入正则化与标签平滑缓解过拟合。在ARIQA数据集上的实验表明,所提方法在所有变体(TransformAR、TransformAR-KD、TransformAR-KD+)中均优于现有最先进方法。

原文摘要 · Abstract (English)

Augmented Reality (AR) is a major immersive media technology that enriches our perception of reality by overlaying digital content (the foreground) onto physical environments (the background). It has far-reaching applications, from entertainment and gaming to education, healthcare, and industrial training. Nevertheless, challenges such as visual confusion and classical distortions can result in user discomfort when using the technology. Evaluating AR quality of experience becomes essential to measure user satisfaction and engagement, facilitating the refinement necessary for creating immersive and robust experiences. Though, the scarcity of data and the distinctive characteristics of AR technology render the development of effective quality assessment metrics challenging. This paper presents a deep learning-based objective metric designed specifically for assessing image quality for AR scenarios. The approach entails four key steps, (1) fine-tuning a self-supervised pre-trained vision transformer to extract prominent features from reference images and distilling this knowledge to improve representations of distorted images, (2) quantifying distortions by computing shift representations, (3) employing cross-attention-based decoders to capture perceptual quality features, and (4) integrating regularization techniques and label smoothing to address the overfitting problem. To validate the proposed approach, we conduct extensive experiments on the ARIQA dataset. The results showcase the superior performance of our proposed approach across all model variants, namely TransformAR, TransformAR-KD, and TransformAR-KD+ in comparison to existing state-of-the-art methods.

AR质量评估知识蒸馏视觉感知图像表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。