通过深度引导注意力提升图像质量评估的泛化能力
DGIQA: Depth-guided Feature Attention and Refinement for Generalizable Image Quality Assessment
- 引入深度引导跨注意力机制,融合场景深度与空间特征
- 在合成与真实数据集上均达当前最优,尤其擅长自然失真评估
- 适合需要泛化能力强的图像质量评估场景
无参考图像质量评估(NR-IQA)长期面临难以泛化到未见自然失真的挑战。为此,本文提出深度引导交叉注意力与精炼机制(Depth-CAR),将场景深度和空间特征提炼为结构感知表示,融入物体显著性与相对对比度知识,增强特征判别力。同时引入Transformer-CNN桥接(TCB)结构,融合Transformer骨干网络的全局上下文依赖与层级CNN捕捉的局部空间特征。DGIQA模型采用多模态注意力投影函数实现关键特征选择,提升训练与推理效率。实验表明,该模型在合成与真实基准数据集上均达到最先进水平,尤其在跨数据集评估及低光、雾霾、镜头眩光等自然失真场景中表现优异。
原文摘要 · Abstract (English)
A long-held challenge in no-reference image quality assessment (NR-IQA) learning from human subjective perception is the lack of objective generalization to unseen natural distortions. To address this, we integrate a novel Depth-Guided cross-attention and refinement (Depth-CAR) mechanism, which distills scene depth and spatial features into a structure-aware representation for improved NR-IQA. This brings in the knowledge of object saliency and relative contrast of the scene for more discriminative feature learning. Additionally, we introduce the idea of TCB (Transformer-CNN Bridge) to fuse high-level global contextual dependencies from a transformer backbone with local spatial features captured by a set of hierarchical CNN (convolutional neural network) layers. We implement TCB and Depth-CAR as multimodal attention-based projection functions to select the most informative features, which also improve training time and inference efficiency. Experimental results demonstrate that our proposed DGIQA model achieves state-of-the-art (SOTA) performance on both synthetic and authentic benchmark datasets. More importantly, DGIQA outperforms SOTA models on cross-dataset evaluations as well as in assessing natural image distortions such as low-light effects, hazy conditions, and lens flares.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。